Skip to main content
QUICK REVIEW

[论文解读] Multi-Receiver Online Bayesian Persuasion

Matteo Castiglioni, Alberto Marchesi|arXiv (Cornell University)|Jun 11, 2021
Advanced Bandit Algorithms Research被引用 6
一句话总结

本文提出了首个针对无外部性、二元行动的多接收者在线贝叶斯说服框架。针对子模性发送方效用函数,提出了一种基于在线梯度下降方案与近似投影预言机的多项式时间无-(1−1/e)-后悔算法;同时证明,在相同约束下,对于超模或匿名效用函数,不存在多项式时间无-α-后悔算法。

ABSTRACT

Bayesian persuasion studies how an informed sender should partially disclose information to influence the behavior of a self-interested receiver. Classical models make the stringent assumption that the sender knows the receiver's utility. This can be relaxed by considering an online learning framework in which the sender repeatedly faces a receiver of an unknown, adversarially selected type. We study, for the first time, an online Bayesian persuasion setting with multiple receivers. We focus on the case with no externalities and binary actions, as customary in offline models. Our goal is to design no-regret algorithms for the sender with polynomial per-iteration running time. First, we prove a negative result: for any $0 < α\leq 1$, there is no polynomial-time no-$α$-regret algorithm when the sender's utility function is supermodular or anonymous. Then, we focus on the case of submodular sender's utility functions and we show that, in this case, it is possible to design a polynomial-time no-$(1 - \frac{1}{e})$-regret algorithm. To do so, we introduce a general online gradient descent scheme to handle online learning problems with a finite number of possible loss functions. This requires the existence of an approximate projection oracle. We show that, in our setting, there exists one such projection oracle which can be implemented in polynomial time.

研究动机与目标

  • 将在线贝叶斯说服扩展至对抗性选择接收者类型的多接收者场景。
  • 为具有多项式每轮运行时间的发送方设计无-α-后悔算法。
  • 分析在不同发送方效用函数类别(超模、子模、匿名)下此类算法的计算可行性。
  • 证明由直接信号配置诱导的拟阵结构存在多项式时间近似投影预言机。
  • 刻画多接收者说服中在线学习的根本极限,对超模与匿名效用函数得出负面结果。

提出的方法

  • 利用直接信号配置上的拟阵结构形式化多接收者说服问题,其中每个基对应一个有效的信号方案。
  • 将发送方的期望效用表示为拟阵基上的非减子模函数与线性函数之和。
  • 应用在线梯度下降方案以在有限个损失函数存在下最小化后悔,依赖于近似投影预言机。
  • 提出一种新颖的近似投影预言机,可在拟阵结构上以多项式时间计算,从而实现高效的在线更新。
  • 采用 Sviridenko 等人(2017)的算法,在拟阵约束下实现子模效用最大化问题的期望 (1−1/e)-近似。
  • 采用完整信息反馈机制,即发送方在每次迭代后可观察到每个接收者的类型,从而实现自适应信号策略。

实验结果

研究问题

  • RQ1能否为多接收者在线贝叶斯说服问题设计出每轮时间复杂度为多项式的无-α-后悔算法?
  • RQ2当发送方效用为超模或匿名时,此类算法的计算限制是什么?
  • RQ3当发送方效用为子模时,是否可实现常数因子后悔保证(具体为 1−1/e)?
  • RQ4在多接收者说服中,由直接信号配置诱导的拟阵结构是否存在高效近似投影预言机?
  • RQ5在线梯度下降能否在此设置下有效适应具有有限个损失函数的在线学习问题?

主要发现

  • 对于任意 0 < α ≤ 1,当发送方效用为超模或匿名时,不存在多项式时间无-α-后悔算法。
  • 当发送方效用为子模时,存在多项式时间无-(1−1/e)-后悔算法,可实现对最优事后信号方案的 (1−1/e)-近似。
  • 拟阵结构上存在高效近似投影预言机,使得在线梯度下降在此场景中得以应用。
  • 该算法在问题实例规模与逆精度参数的多项式时间内,以高概率实现其后悔保证。
  • 该结果在完整信息反馈与固定接收者类型数量下成立,确保了计算可处理性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。