[論文レビュー] Should ChatGPT and Bard Share Revenue with Their Data Providers? A New Business Model for the AI Era
この論文は、ChatGPT や Bard などの大規模なAIモデルが、プロンプトベースのスコアリングシステムに基づいて、訓練データ提供者と収益を共有する、革新的な収益分配モデルを提案する。テキスト分類および類似度モデルを用いてデータの関与度を測定することで、公平でスケーラブルな収益分配が可能となり、敵対的だったデータ関係を協働的で実利的なAIエコシステムに変革する。
With various AI tools such as ChatGPT becoming increasingly popular, we are entering a true AI era. We can foresee that exceptional AI tools will soon reap considerable profits. A crucial question arise: should AI tools share revenue with their training data providers in additional to traditional stakeholders and shareholders? The answer is Yes. Large AI tools, such as large language models, always require more and better quality data to continuously improve, but current copyright laws limit their access to various types of data. Sharing revenue between AI tools and their data providers could transform the current hostile zero-sum game relationship between AI tools and a majority of copyrighted data owners into a collaborative and mutually beneficial one, which is necessary to facilitate the development of a virtuous cycle among AI tools, their users and data providers that drives forward AI technology and builds a healthy AI ecosystem. However, current revenue-sharing business models do not work for AI tools in the forthcoming AI era, since the most widely used metrics for website-based traffic and action, such as clicks, will be replaced by new metrics such as prompts and cost per prompt for generative AI tools. A completely new revenue-sharing business model, which must be almost independent of AI tools and be easily explained to data providers, needs to establish a prompt-based scoring system to measure data engagement of each data provider. This paper systematically discusses how to build such a scoring system for all data providers for AI tools based on classification and content similarity models, and outlines the requirements for AI tools or third parties to build it. Sharing revenue with data providers using such a scoring system would encourage more data owners to participate in the revenue-sharing program. This will be a utilitarian AI era where all parties benefit.
研究の動機と目的
- AI訓練における倫理的・経済的不均衡を是正するため、AIツールがデータ提供者と収益を共有することを提案すること。
- 生成的AIの文脈において、クリック数など古くなった指標に依存する従来の収益分配モデルの限界を克服すること。
- プロンプトとの相互作用に基づいてデータ提供者の関与度を測定する、スケーラブルで技術的に実現可能かつ説明可能なスコアリングシステムを設計すること。
- 大規模言語モデルにとどまらず、画像生成や医療AIなどの他のAIアプリケーションへもモデルを拡張すること。
- AI開発者、ユーザー、データ提供者の間でインcentiveを一致させることで、イノベーションの好循環を促進すること。
提案手法
- ユーザーのプロンプト/生成物と各提供者の訓練データとの間のテキスト類似度を測定することで、データ関与度を定量化するプロンプトベースのスコアリングシステムを開発する。
- 汎用的またはモデル固有のテキスト埋め込み技術を用いて、ドキュメントをベクトルに変換し、類似度計算を実行する。
- 分類モデルを監視学習で適用し、提供者ごとにデータをグループ化することで、スケーラブルな提供者単位のスコアリングを可能にする。
- 複数のプロンプトにおける平均類似度スコアを計算して、各データ提供者ごとの累積関与スコアを生成する。
- 生産規模のAIシステムをサポートするため、最適化された計算複雑度を備えたリアルタイムスコアリングを実装する。
- 画像やその他のデータタイプのためのモダリティ特化スコアリングシステムを開発することで、マルチモーダルAIへのフレームワークの拡張を図る。

実験結果
リサーチクエスチョン
- RQ1大規模言語モデル向けに、データ提供者の貢献を考慮した公平でスケーラブルな収益分配メカニズムをどのように設計できるか。
- RQ2生成的AIシステムにおいて、従来のウェブトラフィック指標に代わる、データ関与度を測定するための指標と技術的要素は何か。
- RQ3ユーザー個別の追跡や中央集権的データ保存に依存せずに、プロンプト相互作用に基づいてデータ提供者をどのようにスコアリングできるか。
- RQ4このフレームワークは、テキストから画像への変換生成ツールや医療AIシステムなどの非LLM AIツールに対しても同様に適応可能か。
- RQ5このようなシステムを大規模に実装する際の技術的・経済的制約は何か。
主な発見
- テキスト類似度と分類モデルを用いたプロンプトベースのスコアリングシステムは、生成的AIアプリケーションにおけるデータ提供者の関与度を効果的に測定できる。
- テキスト埋め込みの使用により、訓練データの出典が完全に開示されていなくても、データ貢献の定量的評価が可能になる。
- 提案されたモデルはリアルタイムスコアリングをサポートでき、画像や医療データを含むマルチモーダルデータへも拡張可能である。
- 画像ベースのAIでは、画像分類および類似度モデルを用いた類似のシステムを構築でき、個々の芸術的作品の関与スコアを算出可能である。
- 訓練データセットにまだ含まれていないデータ提供者に対しても一時的なスコアリングが可能であり、収益分配プログラムへの早期参加を可能にする。
- このモデルは技術的に実現可能であり、Pile や LIOAN-5B データセットなどの既存のNLPおよびコンピュータビジョン技術を用いてベンチマーク化可能である。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。