Skip to main content
QUICK REVIEW

[論文レビュー] Competing Bandits: The Perils of Exploration Under Competition

Guy Aridor, Yishay Mansour|arXiv (Cornell University)|Jul 20, 2020
Auction Theory and Applications被引用数 7
ひとこと要約

本稿は、プラットフォーム間の競争がマルチアームド・バンディット問題における探索戦略の選択に与える影響を検討する。競争が激しいと、企業は非探索的で利益志向の強いアルゴリズムを採用する傾向にあり、結果としてユーザーの福祉が低下する。一方、先行者優位性や非合理的なユーザー選択によって競争を緩和すると、より良い探索が促進され、ユーザー福祉が向上する。

ABSTRACT

Most online platforms strive to learn from interactions with users, and many engage in exploration: making potentially suboptimal choices for the sake of acquiring new information. We study the interplay between exploration and competition: how such platforms balance the exploration for learning and the competition for users. Here users play three distinct roles: they are customers that generate revenue, they are sources of data for learning, and they are self-interested agents which choose among the competing platforms. We consider a stylized duopoly model in which two firms face the same multi-armed bandit problem. Users arrive one by one and choose between the two firms, so that each firm makes progress on its bandit problem only if it is chosen. Through a mix of theoretical results and numerical simulations, we study whether and to what extent competition incentivizes the adoption of better bandit algorithms, and whether it leads to welfare increases for users. We find that stark competition induces firms to commit to a "greedy" bandit algorithm that leads to low welfare. However, weakening competition by providing firms with some "free" users incentivizes better exploration strategies and increases welfare. We investigate two channels for weakening the competition: relaxing the rationality of users and giving one firm a first-mover advantage. Our findings are closely related to the "competition vs. innovation" relationship, and elucidate the first-mover advantage in the digital economy.

研究の動機と目的

  • プラットフォーム間の競争が、ユーザー学習のためのバンディットアルゴリズム採用に与える影響を分析すること。
  • 競争が短期的評判懸念のため、より良い探索を促進するのか、それとも福祉の損失を引き起こすのかを調査すること。
  • ユーザー行動および企業の位置付け(例:先行者優位性)が、探索と搾取のバランスに与える影響を明らかにすること。
  • 不良な探索がユーザーの離脱を引き起こし、学習能力のさらなる低下を招くフィードバックループをモデル化すること。
  • 競争的ダイナミクス下でのデジタル市場における参入障壁としてのデータ格差の役割を検討すること。

提案手法

  • 同じマルチアームド・バンディット問題に直面する二社のスタイリゼッド・ダポリーオモデルを用いる。ユーザーは評判に基づいて選択する。
  • 応答関数(HardMax、SoftMax、HardMax&Random)を用いたベイジアン選択モデルを導入し、ユーザーの合理的さの違いを形式化する。
  • 複数のMABインスタンス(例:Needle-in-Haystack、Balanced、Sparse)に対して数値シミュレーションを実施し、競争下でのアルゴリズム性能を評価する。
  • レピュテーションスコアとレピュテーション差の時間的推移を分析し、アルゴリズムの成果を評価。分布的洞察を得るため、カーネル密度推定を用いる。
  • ユーザーの合理的さを緩和し、先行者優位性を導入することで、アルゴリズム採用とユーザー福祉に与える影響を評価する。
  • 理論的分析により、完全な競争下では、より良いアルゴリズムが存在しても企業がグリーディ戦略に収束することを示す。

実験結果

リサーチクエスチョン

  • RQ1激しい競争は、企業がより良い探索アルゴリズムを採用するのを促進するのか、それとも非探索的で利益志向の強い戦略を採用させるのか?
  • RQ2ユーザー選択ルール(例:合理的 vs. ストキャスティック)が、ダポリーオ・バンディット設定における均衡結果に与える影響は何か?
  • RQ3先行者優位性が探索のインcentiveをどの程度変化させ、ユーザー福祉を向上させるのか?
  • RQ4データやレピュテーション格差が、明示的なモデル化がなくても内生的ネットワーク効果を生じさせる可能性はあるか?
  • RQ5平均レピュテーショントレースが、競争的ダポリーオにおいてアルゴリズムのパフォーマンスを予測できないのはなぜか?

主な発見

  • 激しい競争は、探索を回避するグリーディ・バンディットアルゴリズムの採用を企業に促し、結果としてユーザー福祉が低下する。
  • 先行者優位性や非合理的なユーザー行動によって競争を緩和すると、企業はより良い探索戦略を採用するようになる。
  • BayesGreedyのようなアルゴリズムのレピュテーション分布は二峰性を示す——つまり、トムソン・サンプリングよりわずかに優れるか、著しく劣る——このため、平均レピュテーションは性能予測に不適切である。
  • トムソン・サンプリングとBayesGreedyの間のレピュテーション差は右に歪んでいる。これは、競争的結果において平均が中心傾向の信頼できる指標ではないことを示している。
  • 初期の小さなデータまたはレピュテーション優位性でさえ、時間経過とともに拡大し、市場シェアの大きな格差を生み出し、内生的ネットワーク効果を生成する。
  • 推論5.2(平均レピュテーショントレースが競争的結果を説明できる)は、報酬ベクトルごとのレピュテーションスコアの複雑で非正規分布性のため、経験的に反証された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。