Skip to main content
QUICK REVIEW

[论文解读] Real World Games Look Like Spinning Tops

Wojciech Marian Czarnecki, Gauthier Gidel|arXiv (Cornell University)|Apr 20, 2020
Artificial Intelligence in Games参考文献 34被引用 11
一句话总结

本文提出,现实世界的游戏表现出一种‘陀螺’几何结构,其纵向轴具有传递性强度,径向轴则具有循环性、非传递性动态。通过纳什聚类和对九款游戏的实证分析,本文证明该几何结构自然涌现,并解释了为何在这些游戏中,基于种群的训练对收敛至关重要,尤其是在种群规模超过临界阈值时。

ABSTRACT

This paper investigates the geometrical properties of real world games (e.g. Tic-Tac-Toe, Go, StarCraft II). We hypothesise that their geometrical structure resemble a spinning top, with the upright axis representing transitive strength, and the radial axis, which corresponds to the number of cycles that exist at a particular transitive strength, representing the non-transitive dimension. We prove the existence of this geometry for a wide class of real world games, exposing their temporal nature. Additionally, we show that this unique structure also has consequences for learning - it clarifies why populations of strategies are necessary for training of agents, and how population size relates to the structure of the game. Finally, we empirically validate these claims by using a selection of nine real world two-player zero-sum symmetric games, showing 1) the spinning top structure is revealed and can be easily re-constructed by using a new method of Nash clustering to measure the interaction between transitive and cyclical strategy behaviour, and 2) the effect that population size has on the convergence in these games.

研究动机与目标

  • 识别人类参与的现实世界游戏(如围棋、星际争霸II和DOTA 2)中普遍存在的几何结构。
  • 形式化‘技能之戏’假说,提出这些游戏具有结合传递性强度与循环动态的陀螺几何结构。
  • 从理论和实证两方面验证现实世界游戏中长循环的存在,并将其与学习动态联系起来。
  • 解释为何在这些游戏中,基于种群的训练方法对收敛至关重要,尤其从纳什聚类覆盖的视角进行分析。
  • 区分现实世界游戏与缺乏此类几何结构的经典博弈论游戏(如Blotto游戏)

提出的方法

  • 将现实世界游戏建模为陀螺结构,其中纵向轴表示传递性强度,径向轴表示非传递性循环。
  • 引入纳什聚类方法,用于度量传递性与循环性策略行为之间的相互作用,从而实现陀螺结构的重建。
  • 通过代理种群的实证采样来近似现实世界游戏,并在不同种群规模下评估收敛性。
  • 应用基于悲观oracle的强化学习算法,选择能够击败当前种群的策略,模拟贪婪强化学习。
  • 理论分析证明n步游戏中存在极长的循环,包括所有研究的现实世界游戏。
  • 对比现实世界游戏与非现实世界游戏(如Blotto游戏)的结果,验证‘技能之戏’的几何结构独特性。

实验结果

研究问题

  • RQ1现实世界游戏(如井字棋、围棋和星际争霸II)是否表现出结合传递性强度与循环动态的陀螺几何结构?
  • RQ2能否通过纳什聚类与代理种群采样,实证重建陀螺几何结构?
  • RQ3种群规模与现实世界游戏中学习收敛性之间存在何种关系?
  • RQ4为何基于种群的训练方法在现实世界游戏中成功,却在某些经典博弈论游戏中失败?
  • RQ5在弱策略区域中存在长循环时,为何需要多样化的策略种群?

主要发现

  • 在测试的九款现实世界双人零和对称游戏中(包括3×3围棋、井字棋和星际争霸II),均实证观察到陀螺几何结构。
  • 纳什聚类大小在中等传递性强度处达到峰值,并在高、低技能水平处减小,证实了该几何结构的陀螺形态。
  • 对于现实世界游戏,使用小规模种群训练会导致持续循环;只有当种群规模超过临界阈值时,收敛才会发生。
  • 相比之下,非现实世界游戏(如Blotto游戏)在高技能水平下纳什聚类大小持续增加,且不表现出陀螺结构。
  • Elo游戏(纯粹传递性游戏)即使种群规模为1也能收敛,证实单调进步并不需要种群多样性。
  • 理论结果证明n步游戏包含极长的循环,解释了在学习过程中需要多样化策略种群的必要性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。