Skip to main content
QUICK REVIEW

[论文解读] Policy Regret in Repeated Games

Raman Arora, Michael Dinitz|arXiv (Cornell University)|Nov 9, 2018
Advanced Bandit Algorithms Research被引用 5
一句话总结

本文引入策略后悔(policy regret)作为对抗自适应对手的重复博弈中的性能度量,表明无策略后悔算法会收敛到一种新的均衡概念——策略均衡(policy equilibrium)。论文证明,在自利博弈设定下,无外部后悔(no-external regret)与无策略后悔是相容的,且粗略相关均衡(coarse correlated equilibria,由无外部后悔产生)是策略均衡的真子集,因此策略后悔在博弈论在线学习中是一种更强且更精细的解概念。

ABSTRACT

The notion of \emph{policy regret} in online learning is a well defined? performance measure for the common scenario of adaptive adversaries, which more traditional quantities such as external regret do not take into account. We revisit the notion of policy regret and first show that there are online learning settings in which policy regret and external regret are incompatible: any sequence of play that achieves a favorable regret with respect to one definition must do poorly with respect to the other. We then focus on the game-theoretic setting where the adversary is a self-interested agent. In that setting, we show that external regret and policy regret are not in conflict and, in fact, that a wide class of algorithms can ensure a favorable regret with respect to both definitions, so long as the adversary is also using such an algorithm. We also show that the sequence of play of no-policy regret algorithms converges to a \emph{policy equilibrium}, a new notion of equilibrium that we introduce. Relating this back to external regret, we show that coarse correlated equilibria, which no-external regret players converge to, are a strict subset of policy equilibria. Thus, in game-theoretic settings, every sequence of play with no external regret also admits no policy regret, but the converse does not hold.

研究动机与目标

  • 分析在对抗自适应对手的在线学习中,策略后悔与外部后悔之间的相容性。
  • 定义并表征一种新的均衡概念——策略均衡,作为无策略后悔算法的极限点。
  • 证明在自利博弈设定下,当双方均使用此类算法时,可同时实现无外部后悔与无策略后悔。
  • 建立粗略相关均衡是策略均衡的真子集,从而证明策略后悔作为解概念的优越性。

提出的方法

  • 提出策略后悔作为基准,将在线博弈行为与固定动作策略进行比较,且独立于玩家自身的行为。
  • 引入策略均衡的概念,作为无策略后悔算法生成序列的极限点。
  • 使用马尔可夫链建模来表示记忆受限对手下动作序列的演化过程。
  • 应用小批量简化技术,将无外部后悔算法转换为在有界记忆下实现无策略后悔的算法。
  • 通过策略组合的概率分布分析,证明收敛至策略均衡。
  • 将框架扩展至具有异质记忆边界的多玩家博弈,并在联合动作历史上定义广义马尔可夫过程。

实验结果

研究问题

  • RQ1在具有自适应对手的重复博弈中,无策略后悔与无外部后悔能否共存?
  • RQ2粗略相关均衡与新提出的策略均衡概念之间存在何种关系?
  • RQ3无策略后悔算法生成的动作序列是否收敛至稳定结果?若是,其性质为何?
  • RQ4在博弈论设定下,策略后悔是否是外部后悔的更精细解概念?
  • RQ5能否为具有记忆受限对手的一般多玩家博弈构造无策略后悔算法?

主要发现

  • 在一般在线学习设定中,策略后悔与外部后悔不相容,因为任何在某一指标上实现次线性后悔的算法,可能在另一指标上产生线性后悔。
  • 在自利博弈设定下,当双方均使用确保在两项指标上均实现次线性后悔的算法时,可同时实现无外部后悔与无策略后悔。
  • 无策略后悔算法生成的动作序列收敛至策略均衡,这是一种推广并精炼粗略相关均衡的新解概念。
  • 粗略相关均衡是策略均衡的真子集,即所有无外部后悔的结果都是策略均衡,但反之不成立。
  • 通过一个具有非对称收益的两玩家示例,证明无策略后悔策略所实现的效用可严格高于无外部后悔策略。
  • 本文构造了明确的策略均衡实例,其并非粗略相关均衡,从而验证了严格精炼性质。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。