Skip to main content
QUICK REVIEW

[论文解读] Competing Bandits: Learning under Competition

Yishay Mansour, Aleksandrs Slivkins|arXiv (Cornell University)|Feb 27, 2017
Advanced Bandit Algorithms Research参考文献 27被引用 19
一句话总结

本文研究了两个多臂赌博机算法在用户之间竞争的场景,用户根据期望效用进行选择。研究发现,在高理性度(HardMax)下,竞争会抑制探索行为,导致采用贪婪策略;但在中等理性度(SoftMax)下,更优的算法反而更可能被采纳——揭示了竞争性与创新采纳之间存在倒U型关系。

ABSTRACT

Most modern systems strive to learn from interactions with users, and many engage in exploration: making potentially suboptimal choices for the sake of acquiring new information. We initiate a study of the interplay between exploration and competition--how such systems balance the exploration for learning and the competition for users. Here the users play three distinct roles: they are customers that generate revenue, they are sources of data for learning, and they are self-interested agents which choose among the competing systems. In our model, we consider competition between two multi-armed bandit algorithms faced with the same bandit instance. Users arrive one by one and choose among the two algorithms, so that each algorithm makes progress if and only if it is chosen. We ask whether and to what extent competition incentivizes the adoption of better bandit algorithms. We investigate this issue for several models of user response, as we vary the degree of rationality and competitiveness in the model. Our findings are closely related to the "competition vs. innovation" relationship, a well-studied theme in economics.

研究动机与目标

  • 理解学习系统之间的竞争如何影响更优探索算法的采纳。
  • 在多臂赌博机设置中,建模用户理性、市场竞争力与算法选择之间的相互作用。
  • 分析竞争是否激励采用更优的学习算法,或因缺乏探索而导致次优结果。
  • 研究在不同用户响应模型下,更优算法在均衡中出现的条件。
  • 探讨在用户主导选择的学习系统中,竞争与垄断对社会福利的影响。

提出的方法

  • 建模一个博弈场景:两名主体使用多臂赌博机算法,并竞争吸引按顺序到达的用户。
  • 用户根据响应函数(如 HardMax、SoftMax 或 HardMax&Random)选择主体,以反映其理性程度和决策行为。
  • 每位主体仅能观察选择自己的用户,必须在不观察对方结果的情况下学习自身动作的奖励分布。
  • 分析不同响应函数下的均衡行为,重点关注更优算法(如具有更低贝叶斯遗憾的算法)是否被采纳。
  • 引入 BIR-主导性(Bayesian-Improvement-Rational dominance)概念,以刻画更优算法被偏好的条件。
  • 通过理论分析证明在各种响应函数下(尤其是 SoftMax 和 HardMax&Random)纳什均衡的存在性与唯一性。

实验结果

研究问题

  • RQ1学习系统之间的竞争是否激励采用更优的多臂赌博机算法?
  • RQ2用户理性度如何影响竞争环境下学习算法的均衡选择?
  • RQ3在何种条件下,具有更低贝叶斯遗憾的更优算法会在均衡中占主导地位?
  • RQ4竞争性与学习系统中创新之间的关系如何?是否呈现倒U型模式?
  • RQ5用户响应结构(如 SoftMax 与 HardMax)如何影响学习算法的长期表现与采纳情况?

主要发现

  • 在 HardMax 响应函数(完全理性)下,主导策略为 DynamicGreedy,其避免探索,从而阻碍更优算法的采纳。
  • 在 HardMax&Random 且随机化程度较低(ε₀ 接近 0)时,若某算法 BIR-主导其他算法,可在均衡中被采纳,但收益提升有限。
  • 在 SoftMax 响应下,只要某算法弱 BIR-主导其他算法,更优算法即可被采纳,表明创新激励更强。
  • 随着用户理性度(由 ε₀ 控制)变化,切换至更优算法的边际效用呈倒U型,中等理性水平时达到峰值。
  • 在均匀选择模型(ε₀ = 0.5)下,不存在创新激励,因为用户无论算法质量如何均随机选择。
  • 垄断情形下,社会福利可能优于竞争,因为单一主体有激励最小化遗憾并采纳最优算法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。