Skip to main content
QUICK REVIEW

[论文解读] Competing Bandits: The Perils of Exploration Under Competition

Guy Aridor, Yishay Mansour|arXiv (Cornell University)|Jul 20, 2020
Auction Theory and Applications被引用 7
一句话总结

本文研究平台间竞争如何影响其在多臂赌博机问题中选择探索策略。研究发现,激烈竞争促使企业采用贪婪的、非探索性算法,从而导致用户福利下降;而通过先行者优势或用户非理性选择减少竞争,则能促进更好的探索行为并提升用户福利。

ABSTRACT

Most online platforms strive to learn from interactions with users, and many engage in exploration: making potentially suboptimal choices for the sake of acquiring new information. We study the interplay between exploration and competition: how such platforms balance the exploration for learning and the competition for users. Here users play three distinct roles: they are customers that generate revenue, they are sources of data for learning, and they are self-interested agents which choose among the competing platforms. We consider a stylized duopoly model in which two firms face the same multi-armed bandit problem. Users arrive one by one and choose between the two firms, so that each firm makes progress on its bandit problem only if it is chosen. Through a mix of theoretical results and numerical simulations, we study whether and to what extent competition incentivizes the adoption of better bandit algorithms, and whether it leads to welfare increases for users. We find that stark competition induces firms to commit to a "greedy" bandit algorithm that leads to low welfare. However, weakening competition by providing firms with some "free" users incentivizes better exploration strategies and increases welfare. We investigate two channels for weakening the competition: relaxing the rationality of users and giving one firm a first-mover advantage. Our findings are closely related to the "competition vs. innovation" relationship, and elucidate the first-mover advantage in the digital economy.

研究动机与目标

  • 分析平台之间的竞争如何影响其在用户学习中采用赌博机算法的决策。
  • 研究竞争是否激励企业采取更好的探索策略,或因短期声誉担忧而导致福利损失。
  • 探讨用户行为及企业定位(如先行者优势)如何影响探索与利用之间的平衡。
  • 建立反馈回路模型,说明不良探索如何导致用户流失并进一步削弱学习能力。
  • 考察数据差异在竞争动态下的数字市场中作为进入壁垒的作用。

提出的方法

  • 采用简化的双寡头模型,其中两家公司面临相同的多臂赌博机问题,并根据声誉竞争用户。
  • 引入贝叶斯选择模型与响应函数(HardMax、SoftMax、HardMax&Random),以形式化不同理性程度下的用户选择行为。
  • 通过多个MAB实例(如Needle-in-Haystack、Balanced和Sparse)的数值模拟,评估竞争环境下算法的表现。
  • 分析随时间推移的声誉得分及其差异,使用核密度估计法获取分布特征的洞察。
  • 评估放松用户理性假设及引入先行者优势对算法采纳与用户福利的影响。
  • 通过理论分析表明,在完全竞争条件下,企业即使存在更优算法,仍会收敛至贪婪策略。

实验结果

研究问题

  • RQ1激烈的竞争是否激励企业采用更好的探索算法,还是迫使其转向贪婪的、非探索性策略?
  • RQ2用户选择规则(如理性选择与随机选择)如何影响双寡头赌博机设定下的均衡结果?
  • RQ3先行者优势在多大程度上改变探索激励并改善用户福利?
  • RQ4数据或声誉差异是否可能在未明确建模的情况下引发内生网络效应?
  • RQ5为何均值声誉轨迹无法准确预测竞争双寡头中的算法表现?

主要发现

  • 激烈的竞争导致企业采用避免探索的贪婪赌博机算法,从而降低用户福利。
  • 通过先行者优势或用户行为非理性程度降低竞争,可促使企业采纳更优的探索策略。
  • BayesGreedy等算法的声誉分布呈双峰特征——其表现要么略优于Thompson Sampling,要么显著劣于后者,因此均值声誉无法充分预测其表现。
  • Thompson Sampling与BayesGreedy之间的声誉差异呈右偏分布,表明在竞争结果中,均值并非可靠的集中趋势度量。
  • 即使初始数据或声誉优势微小,也会随时间被放大,导致市场份额的巨大差异,并内生性地产生网络效应。
  • 推测5.2(即均值声誉轨迹可解释竞争结果)因声誉得分在奖励向量上的分布复杂且非正态,被实证所证伪。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。