Skip to main content
QUICK REVIEW

[论文解读] Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum Games

Xiangyu Liu, Hangtian Jia|arXiv (Cornell University)|Jun 9, 2021
Experimental Behavioral Economics Studies参考文献 3被引用 9
一句话总结

本文提出了一种统一框架,用于在零和马尔可夫游戏中对开放式学习中的行为多样性(BD)和响应多样性(RD)进行测量与促进。通过将BD定义为状态-动作空间中占据测度的差异,将RD通过投影到游戏景观凸包上的几何距离来定义,并引入下界以实现实际优化,该方法增强了策略群体的多样性,从而在矩阵博弈中实现更低的可被利用性与更高的群体有效性,在Google Research Football中也表现出更优性能。

ABSTRACT

Measuring and promoting policy diversity is critical for solving games with strong non-transitive dynamics where strategic cycles exist, and there is no consistent winner (e.g., Rock-Paper-Scissors). With that in mind, maintaining a pool of diverse policies via open-ended learning is an attractive solution, which can generate auto-curricula to avoid being exploited. However, in conventional open-ended learning algorithms, there are no widely accepted definitions for diversity, making it hard to construct and evaluate the diverse policies. In this work, we summarize previous concepts of diversity and work towards offering a unified measure of diversity in multi-agent open-ended learning to include all elements in Markov games, based on both Behavioral Diversity (BD) and Response Diversity (RD). At the trajectory distribution level, we re-define BD in the state-action space as the discrepancies of occupancy measures. For the reward dynamics, we propose RD to characterize diversity through the responses of policies when encountering different opponents. We also show that many current diversity measures fall in one of the categories of BD or RD but not both. With this unified diversity measure, we design the corresponding diversity-promoting objective and population effectivity when seeking the best responses in open-ended learning. We validate our methods in both relatively simple games like matrix game, non-transitive mixture model, and the complex extit{Google Research Football} environment. The population found by our methods reveals the lowest exploitability, highest population effectivity in matrix game and non-transitive mixture model, as well as the largest goal difference when interacting with opponents of various levels in extit{Google Research Football}.

研究动机与目标

  • 解决开放式学习中零和博弈缺乏一致多样性定义的问题。
  • 将行为多样性与响应多样性统一为一个理论基础坚实的单一度量,适用于马尔可夫游戏。
  • 通过多智能体训练中的多样性促进目标,提升群体有效性并降低可被利用性。
  • 提出群体有效性作为评估策略群体强度的更公平替代指标,以替代可被利用性。
  • 在矩阵博弈、非传递性混合模型以及Google Research Football环境中对方法进行验证。

提出的方法

  • 将行为多样性(BD)定义为状态-动作空间中策略占据测度之间的f-散度。
  • 将响应多样性(RD)定义为策略响应向量到游戏景观凸包的几何距离,并引入下界以实现实际优化。
  • 制定一个统一的多样性目标,结合BD与RD,并使用可学习权重λ₁与λ₂来平衡两个分量。
  • 利用元博弈中的纳什均衡计算对手混合策略,以支持最佳响应计算。
  • 采用强化学习方法或简化的梯度优化方法(例如通过定理1与定理2)更新策略,以实现多样化且高效率的响应。
  • 实现两种算法变体:一种用于矩阵博弈(算法3),另一种用于微分博弈(算法4),并为BD与RD分别设计独立的优化目标。

实验结果

研究问题

  • RQ1如何在零和博弈的多智能体开放式学习中,将行为多样性与响应多样性正式统一为单一度量?
  • RQ2在非传递性博弈中,联合优化BD与RD对群体可被利用性与有效性的影响如何?
  • RQ3所提出的群体有效性度量与可被利用性相比,在评估策略群体强度方面表现如何?
  • RQ4该统一多样性度量在Google Research Football等复杂环境中在多大程度上提升了性能?
  • RQ5超参数λ₁与λ₂控制下的行为多样性与响应多样性之间的最优平衡是什么?

主要发现

  • 在矩阵博弈中,所提方法相比基线方法实现了最低的可被利用性与最高的群体有效性,其中PSRO结合BD&RD优于仅使用RD的PSRO。
  • 在非传递性混合模型中,统一多样性方法生成的策略群体比仅依赖单一多样性类型的策略更有效且更具鲁棒性。
  • 在Google Research Football中,该方法在面对不同技能水平的对手时实现了最大的进球差,展现出卓越的泛化能力与适应性。
  • 消融实验表明,将λ₂(响应多样性权重)设为0.5时性能最佳,表明两种多样性类型对性能有均衡贡献。
  • 群体有效性度量被证明是比可被利用性更公平、更具信息量的评估指标,尤其在非传递性设置中表现更优。
  • 统一框架成功捕捉了策略行为与响应动态,并在多个环境中获得了理论与实证验证。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。