Skip to main content
QUICK REVIEW

[论文解读] Multi-Armed Bandit Mechanisms for Multi-Slot Sponsored Search Auctions

Akash Das Sarma, Sujit Gujar|arXiv (Cornell University)|Jan 9, 2010
Advanced Bandit Algorithms Research参考文献 7被引用 3
一句话总结

本文为多时段赞助搜索拍卖中的真实多臂赌博(MAB)机制提供了理论框架,其中点击率(CTRs)初始未知。研究发现,在无约束CTRs下,真实机制的遗憾为O(T);而在具有结构化CTRs(如点击优先级、粗略预估或可分CTRs)时,弱点态单调性和弱分离分配是确保真实性的充要条件,遗憾被限制在O(T^{2/3})以内。

ABSTRACT

In pay-per click sponsored search auctions which are currently extensively used by search engines, the auction for a keyword involves a certain number of advertisers (say k) competing for available slots (say m) to display their ads. This auction is typically conducted for a number of rounds (say T). There are click probabilities mu_ij associated with each agent-slot pairs. The goal of the search engine is to maximize social welfare of the advertisers, that is, the sum of values of the advertisers. The search engine does not know the true values advertisers have for a click to their respective ads and also does not know the click probabilities mu_ij s. A key problem for the search engine therefore is to learn these click probabilities during the T rounds of the auction and also to ensure that the auction mechanism is truthful. Mechanisms for addressing such learning and incentives issues have recently been introduced and are aptly referred to as multi-armed-bandit (MAB) mechanisms. When m = 1, characterizations for truthful MAB mechanisms are available in the literature and it has been shown that the regret for such mechanisms will be O(T^{2/3}). In this paper, we seek to derive a characterization in the realistic but non-trivial general case when m > 1 and obtain several interesting results.

研究动机与目标

  • 设计适用于初始CTRs未知的多时段赞助搜索拍卖的诚实多臂赌博机制。
  • 刻画在存在策略性出价者和不完全CTRs信息条件下,MAB机制保持诚实的条件。
  • 分析在各种CTRs假设(包括无约束、优先级排序、预估和可分CTRs)下,诚实机制的遗憾边界。
  • 将先前针对单时段MAB机制的研究成果扩展至更具现实意义的多时段设置,涵盖多个竞争广告商和多个广告位。
  • 通过模拟评估所提出诚实MAB机制的性能,重点关注平均遗憾和最坏情况下的遗憾。

提出的方法

  • 提出一种随机多轮拍卖模型,广告商出价其真实估值,搜索引擎通过T轮重复拍卖学习CTRs。
  • 引入弱点态单调性和弱分离分配的概念,作为在结构化CTRs假设下MAB机制保持诚实的关键条件。
  • 基于基于估计CTRs(μ′)的期望效用最大化,推导出支付规则,确保在期望意义上实现诚实出价。
  • 在初始探索阶段(T^{2/3}轮)采用轮换分配策略,以估计CTRs,随后进入利用阶段。
  • 通过比较实际社会福利与在完全掌握CTRs信息下可实现的最优福利,进行遗憾分析。
  • 在k=4个广告商和m=2个广告位的设置下进行模拟,生成100T个随机实例,以估计平均遗憾和最坏情况下的遗憾。

实验结果

研究问题

  • RQ1在CTRs未知的情况下,何种分配规则条件可确保多时段赞助搜索拍卖中机制的诚实性?
  • RQ2在无约束CTRs下,诚实MAB机制的遗憾如何随拍卖轮数T变化?
  • RQ3当CTRs满足如优先级或可分性等结构性特征时,是否较弱的单调性条件(如弱点态单调性)足以保证诚实性?
  • RQ4在现实CTRs假设下,多时段设置中诚实MAB机制可达到的遗憾边界是多少?
  • RQ5所提出的诚实MAB机制在模拟中,其平均遗憾和最坏情况遗憾行为与理论边界相比如何?

主要发现

  • 在无约束CTRs下,任何诚实MAB机制都必须满足强点态单调性,且此类机制的遗憾为O(T)。
  • 当CTRs满足点击优先级性质时,弱点态单调性是MAB机制保持诚实的充要条件。
  • 在CTRs具有粗略预估的情况下,弱点态单调性和弱分离分配共同构成诚实MAB机制的充要条件。
  • 对于可分CTRs(μ_ij = α_i * β_j),本文给出了在期望意义上诚实的MAB机制的完整刻画。
  • 模拟结果表明,最坏情况遗憾的规模为O(T^{2/3}),其近似值为(17/3)T^{2/3},而平均情况遗憾被(1/3)T^{2/3}所限制,与理论边界一致。
  • 实验评估证实,所提出的诚实MAB机制在平均和最坏情况下的遗憾均控制在理论O(T^{2/3})边界之内。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。