Skip to main content
QUICK REVIEW

[论文解读] Chaos, Extremism and Optimism: Volume Analysis of Learning in Games

Yun Kuen Cheung, Georgios Piliouras|arXiv (Cornell University)|May 28, 2020
Reinforcement Learning in Robotics参考文献 20被引用 11
一句话总结

本文通过在对偶(总体收益)空间中引入体积分析,研究博弈中的学习动态,揭示了在零和博弈中,乘法权重更新(MWU)导致体积呈指数级扩张,从而引发李雅普诺夫混沌、极端行为和不可避免性;而在协调博弈中,乐观型乘法权重更新(OMWU)则导致体积收缩,解释了其在零和博弈中的收敛行为和在协调博弈中的混沌行为。

ABSTRACT

We present volume analyses of Multiplicative Weights Updates (MWU) and Optimistic Multiplicative Weights Updates (OMWU) in zero-sum as well as coordination games. Such analyses provide new insights into these game dynamical systems, which seem hard to achieve via the classical techniques within Computer Science and Machine Learning. The first step is to examine these dynamics not in their original space (simplex of actions) but in a dual space (aggregate payoff space of actions). The second step is to explore how the volume of a set of initial conditions evolves over time when it is pushed forward according to the algorithm. This is reminiscent of approaches in Evolutionary Game Theory where replicator dynamics, the continuous-time analogue of MWU, is known to always preserve volume in all games. Interestingly, when we examine discrete-time dynamics, both the choice of the game and the choice of the algorithm play a critical role. So whereas MWU expands volume in zero-sum games and is thus Lyapunov chaotic, we show that OMWU contracts volume, providing an alternative understanding for its known convergent behavior. However, we also prove a no-free-lunch type of theorem, in the sense that when examining coordination games the roles are reversed: OMWU expands volume exponentially fast, whereas MWU contracts. Using these tools, we prove two novel, rather negative properties of MWU in zero-sum games: (1) Extremism: even in games with unique fully mixed Nash equilibrium, the system recurrently gets stuck near pure-strategy profiles, despite them being clearly unstable from game theoretic perspective. (2) Unavoidability: given any set of good points (with your own interpretation of "good"), the system cannot avoid bad points indefinitely.

研究动机与目标

  • 通过一种新颖的体积分析框架,理解MWU与OMWU在零和博弈与协调博弈中行为差异的根源。
  • 识别为何OMWU在零和博弈中收敛而MWU发散,尽管二者更新规则结构相似。
  • 揭示MWU的新负动力学特性,如由体积扩张引发的极端行为与不可避免性。
  • 通过体积变化行为建立零和博弈中MWU与协调博弈中OMWU之间的对偶性。
  • 提供一种原则性且可推广的工具——对偶空间中的体积分析,用于分析超越经典方法的离散时间学习动态。

提出的方法

  • 将学习动态分析从原始单纯形空间(策略组合)转换到对偶空间(总体收益空间),以追踪体积演化。
  • 将体积分析应用于离散时间算法(如MWU与OMWU),测量初始集合的勒贝格测度在迭代过程中的演化。
  • 利用更新映射的雅可比行列式计算体积被积函数,关键表达式涉及曲率项 $ C_{( extbf{A}, extbf{B})}( extbf{p}, extbf{q}) $。
  • 通过雅可比行列式的泰勒展开推导体积变化边界,表明当 $ C_{( extbf{A}, extbf{B})} > 0 $ 时发生扩张,当 $ C_{( extbf{A}, extbf{B})} < 0 $ 时发生收缩。
  • 利用 $ C_{( extbf{A},- extbf{A})} $ 与 $ C_{( extbf{A}, extbf{A})} $ 之间的对偶性,证明零和博弈中的MWU与协调博弈中的OMWU表现出几乎相同的体积动态。
  • 应用关于李雅普诺夫混沌、极端行为与遍历性的结果,推导出在两种情境下不可预测性与不稳定性定理。

实验结果

研究问题

  • RQ1为何MWU在零和博弈中表现出混沌且发散的行为,尽管其时间平均是收敛的?
  • RQ2为何OMWU在零和博弈中实现收敛而MWU不能,尽管其更新规则相似?
  • RQ3除了时间平均之外,导致MWU在零和博弈中不稳定的动力学机制是什么?
  • RQ4对偶空间中体积的扩张或收缩如何与收敛性、遍历性或极端行为相关联?
  • RQ5体积分析能否揭示一个统一原理,解释MWU与OMWU在不同博弈类别中的对比行为?

主要发现

  • 在零和博弈中,MWU在对偶空间中导致体积呈指数级扩张,意味着存在李雅普诺夫混沌和对初始条件的极端敏感性。
  • 在零和博弈中,OMWU导致体积收缩,为其收敛行为提供了几何解释。
  • 在协调博弈中,OMWU导致体积呈指数级扩张,引发李雅普诺夫混沌,而MWU则导致体积收缩。
  • 本文证明,零和博弈中的MWU表现出“极端主义”——即使在不稳定状态下,仍会反复访问接近纯策略组合的区域。
  • MWU还表现出“不可避免性”,即无论“好”点的定义如何,它都无法长期避免不良状态。
  • 由于曲率项 $ C_{( extbf{A}, extbf{B})} $ 的对偶性,零和博弈中MWU的体积动态与协调博弈中OMWU的体积动态几乎完全相同,导致同构的不稳定性模式。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。