Skip to main content
QUICK REVIEW

[论文解读] Mechanism Design with Bandit Feedback.

Kirthevasan Kandasamy, Joseph E. Gonzalez|arXiv (Cornell University)|Apr 19, 2020
Auction Theory and Applications参考文献 44被引用 6
一句话总结

本文提出了一种在代理仅在经历分配后才了解其价值的多轮福利最大化场景下的真实且个体理性机制。该机制引入了一种具有三种相互关联遗憾度量的bandit反馈模型,并建立了Ω(T^{2/3})的下界,通过允许在代理遗憾与卖家遗憾之间权衡的即时算法,实现了该速率,同时保持了真实性。

ABSTRACT

We study a multi-round welfare-maximising mechanism design problem in instances where agents do not know their values. On each round, a mechanism assigns an allocation each to a set of agents and charges them a price; then the agents provide (stochastic) feedback to the mechanism for the allocation they received. This is motivated by applications in cloud markets and online advertising where an agent may know her value for an allocation only after experiencing it. Therefore, the mechanism needs to explore different allocations for each agent, while simultaneously attempting to find the socially optimal set of allocations. Our focus is on truthful and individually rational mechanisms which imitate the classical VCG mechanism in the long run. To that end, we define three notions of regret for the welfare, the individual utilities of each agent and that of the mechanism. We show that these three terms are interdependent via an $\Omega(T^{\frac{2}{3}})$ lower bound for the maximum of these three terms after $T$ rounds of allocations, and describe a family of anytime algorithms which achieve this rate. Our framework provides flexibility to control the pricing scheme so as to trade-off between the agent and seller regrets, and additionally to control the degree of truthfulness and individual rationality.

研究动机与目标

  • 设计一种在代理仅在经历分配后才了解其价值的多轮福利最大化场景下的真实且个体理性的机制。
  • 将反馈过程建模为随机bandit反馈,其中代理在分配后提供反馈。
  • 定义并分析三种相互关联的遗憾项:社会福利、个体代理效用和机制效用。
  • 通过这三项遗憾最大值的下界,建立此类机制性能的根本限制。
  • 开发一族能够实现最优Ω(T^{2/3})遗憾率的即时算法,同时允许在定价、真实性与个体理性之间进行权衡。

提出的方法

  • 该机制在轮次中运行,分配资源并收取价格,随后从代理处接收关于其获得分配质量的随机反馈。
  • 它定义了三种遗憾度量:福利遗憾(与最优社会福利的偏离)、代理效用遗憾(与最优个体效用的偏离)以及机制效用遗憾(与最优机制收入的偏离)。
  • 该框架采用bandit反馈模型,其中代理的价值仅在分配后才被揭示,因此需要探索以学习偏好。
  • 它建立了在T轮后三项遗憾最大值的Ω(T^{2/3})根本下界,表明存在固有的权衡。
  • 所提出的即时算法通过在探索与利用之间取得平衡,同时保持真实性与个体理性,实现了该最优速率。
  • 定价机制具有灵活性,允许在代理遗憾与机制遗憾之间进行权衡,并可调节真实性与个体理性的水平。

实验结果

研究问题

  • RQ1在具有bandit反馈的多轮机制中,福利、代理效用与机制效用遗憾之间存在何种根本性权衡?
  • RQ2是否可以设计出一种真实且个体理性的机制,在此类设置中实现次线性遗憾?若是,其最优速率是多少?
  • RQ3三项遗憾项——福利、代理效用与机制效用——如何相互作用并彼此制约?
  • RQ4在T轮内,这三项遗憾最大值的最紧可能下界是什么?
  • RQ5能否设计出即时算法,以实现最优遗憾速率,同时对真实性与个体理性水平进行控制?

主要发现

  • 本文建立了在T轮后三项遗憾项——福利、代理效用与机制效用——最大值的Ω(T^{2/3})根本下界。
  • 该下界表明,任何机制都无法同时将这三项遗憾低于该速率,凸显了内在的权衡。
  • 作者提出了一族能够实现最优Ω(T^{2/3})遗憾率的即时算法,与下界完全匹配。
  • 该框架允许对定价机制进行灵活控制,以在代理遗憾与机制遗憾之间进行权衡。
  • 这些算法在实现最优遗憾速率的同时,保持了真实性与个体理性。
  • 结果表明,三项遗憾项之间的相互依赖关系无法规避,所提出的算法在理论极限下有效平衡了这些项。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。