Skip to main content
QUICK REVIEW

[论文解读] No-regret learning and mixed Nash equilibria: They do not mix

Lampros Flokas, Emmanouil-Vasileios Vlatakis-Gkaragkounis|arXiv (Cornell University)|Oct 19, 2020
Advanced Bandit Algorithms Research参考文献 59被引用 13
一句话总结

本文证明,基于跟随正则化领导者(FTRL)的无遗憾学习动态无法渐近稳定混合纳什均衡——只有严格(纯)纳什均衡才能成为稳定的极限点。关键结果是,任何非严格的纳什均衡(即至少有一名玩家存在多个最优响应)在FTRL下均无法实现渐近稳定,原因在于FTRL动态在收益空间中具有保体积性质,这阻止了对具有非唯一最优响应的混合策略均衡的收敛。

ABSTRACT

Understanding the behavior of no-regret dynamics in general $N$-player games is a fundamental question in online learning and game theory. A folk result in the field states that, in finite games, the empirical frequency of play under no-regret learning converges to the game's set of coarse correlated equilibria. By contrast, our understanding of how the day-to-day behavior of the dynamics correlates to the game's Nash equilibria is much more limited, and only partial results are known for certain classes of games (such as zero-sum or congestion games). In this paper, we study the dynamics of "follow-the-regularized-leader" (FTRL), arguably the most well-studied class of no-regret dynamics, and we establish a sweeping negative result showing that the notion of mixed Nash equilibrium is antithetical to no-regret learning. Specifically, we show that any Nash equilibrium which is not strict (in that every player has a unique best response) cannot be stable and attracting under the dynamics of FTRL. This result has significant implications for predicting the outcome of a learning process as it shows unequivocally that only strict (and hence, pure) Nash equilibria can emerge as stable limit points thereof.

研究动机与目标

  • 研究无遗憾学习动态(如FTRL)在一般N人博弈中是否能收敛至纳什均衡。
  • 解决无遗憾学习收敛至均衡的民间信念与此类动态下纳什均衡实际稳定性特性之间的脱节问题。
  • 阐明为何粗相关均衡(由无遗憾学习保证存在)并不意味着收敛至纳什均衡,尤其是混合均衡。
  • 在FTRL动态下建立纯均衡与混合均衡稳定性之间的根本性二分。
  • 通过几何与动力系统论证,证明仅严格(纯)纳什均衡可在FTRL下实现渐近稳定。

提出的方法

  • 在收益空间中分析FTRL动态,其在任意基础博弈下均保持勒贝格测度(体积)不变。
  • 利用FTRL动态在收益空间中保体积的性质,当假设非严格混合纳什均衡具有渐近稳定性时,导出矛盾。
  • 应用FTRL在陡峭正则化下对单纯形中面的前向不变性,以限制稳定集的可能位置。
  • 通过反证法证明,任何渐近稳定的集合必须与最小维数的面相交,且该面必须为单点(即纯策略)。
  • 利用博弈单纯形的结构及面的相对内部性质,证明稳定集不可能位于非单点面的内部。
  • 在面的受限博弈中运用李雅普诺夫稳定性与吸引性论证,表明若稳定集位于非单点面的内部,则与保体积性质矛盾。

实验结果

研究问题

  • RQ1在一般N人博弈中,FTRL动态能否使混合纳什均衡实现渐近稳定?
  • RQ2尽管无遗憾学习收敛至粗相关均衡,为何其无法收敛至混合纳什均衡?
  • RQ3FTRL的何种几何或动力学特性阻止其稳定非严格混合纳什均衡?
  • RQ4在无遗憾学习下,纯均衡与混合均衡的稳定性是否存在根本性差异?
  • RQ5FTRL在收益空间中的保体积性质是否可用于排除非严格均衡的渐近稳定性?

主要发现

  • 任何非严格纳什均衡(即至少有一名玩家存在多个最优响应)在FTRL动态下均无法实现渐近稳定。
  • 仅纯纳什均衡可作为FTRL动态的渐近稳定极限点,因为具有非唯一最优响应的混合均衡违反了FTRL在收益空间中的保体积性质。
  • FTRL动态保持了收益空间中集合的勒贝格测度,这是无法稳定非严格混合均衡的根本原因。
  • FTRL下渐近稳定的集合必须与最小维数的面相交,且该面必须为单点,意味着稳定点仅由纯策略构成。
  • 该结果适用于所有具有陡峭正则化的博弈,并适用于整个FTRL动态类,包括乘法权重法与在线梯度下降法。
  • 本文建立了具有广泛影响的负面结论:在FTRL下,混合纳什均衡与渐近稳定性本质上不相容,使其无法成为无遗憾学习的长期结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。