Skip to main content
QUICK REVIEW

[论文解读] Calibration of Shared Equilibria in General Sum Partially Observable Markov Games

Nelson Vadori, Sumitra Ganesh|arXiv (Cornell University)|Jun 23, 2020
Game Theory and Applications参考文献 32被引用 6
一句话总结

本文提出共享均衡(Shared equilibrium)作为在参数共享策略下的通用和部分可观测马尔可夫博弈中对称纯纳什均衡的一种形式,通过自对弈证明其收敛性。提出CAL-SHEQ,一种双强化学习算法,联合训练策略网络与强化学习校准器,以平滑校准智能体类型分布,在匹配现实世界行为目标方面表现优于贝叶斯优化,实现稳定且精确的收敛。

ABSTRACT

Training multi-agent systems (MAS) to achieve realistic equilibria gives us a useful tool to understand and model real-world systems. We consider a general sum partially observable Markov game where agents of different types share a single policy network, conditioned on agent-specific information. This paper aims at i) formally understanding equilibria reached by such agents, and ii) matching emergent phenomena of such equilibria to real-world targets. Parameter sharing with decentralized execution has been introduced as an efficient way to train multiple agents using a single policy network. However, the nature of resulting equilibria reached by such agents has not been yet studied: we introduce the novel concept of Shared equilibrium as a symmetric pure Nash equilibrium of a certain Functional Form Game (FFG) and prove convergence to the latter for a certain class of games using self-play. In addition, it is important that such equilibria satisfy certain constraints so that MAS are calibrated to real world data for practical use: we solve this problem by introducing a novel dual-Reinforcement Learning based approach that fits emergent behaviors of agents in a Shared equilibrium to externally-specified targets, and apply our methods to a n-player market example. We do so by calibrating parameters governing distributions of agent types rather than individual agents, which allows both behavior differentiation among agents and coherent scaling of the shared policy network to multiple agents.

研究动机与目标

  • 正式刻画在通用和部分可观测马尔可夫博弈中,使用共享策略的智能体所达到均衡的游戏理论性质。
  • 解决将涌现的多智能体行为校准至外部指定的现实世界目标(如交易频率或市场份额约束)的挑战。
  • 开发一种高效且可扩展的方法,通过联合学习策略网络与基于强化学习的校准器来优化智能体类型分布,避免重复训练周期。
  • 通过引入双时间尺度随机逼近框架并实现平滑参数更新,确保均衡的稳定收敛。

提出的方法

  • 在函数形式博弈(Functional Form Game, FFG)中引入共享均衡作为对称纯纳什均衡,并在特定游戏条件下通过自对弈证明其收敛性。
  • 提出CAL-SHEQ,一种双强化学习框架:共享策略网络学习均衡,而独立的强化学习校准器则优化控制智能体类型(超类型)分布的参数。
  • 通过调整超类型配置(如连接性、库存容忍度)而非单个智能体,实现系统的一致性扩展与行为差异化。
  • 在校准器中使用加权奖励函数,包含子目标如匹配交易量分位数与外部目标的市场占有率。
  • 采用双时间尺度随机逼近:策略网络快速更新,而校准器慢速更新,以确保在应用新超类型配置前均衡已稳定。
  • 将该方法应用于n方市场模拟,其中不同超类型的商人基于共享策略与目标约束进行交互。

实验结果

研究问题

  • RQ1在通用和部分可观测马尔可夫博弈中,使用共享策略网络的不同类型智能体所达到的均衡具有何种游戏理论性质?
  • RQ2如何将此类多智能体系统中的涌现行为校准至现实世界约束(如交易频率或市场份额)?
  • RQ3能否设计一种联合学习框架,避免昂贵的重复训练周期,同时确保校准均衡的稳定收敛?
  • RQ4与直接调整单个智能体参数相比,校准超类型分布在可扩展性与行为一致性方面有何优势?

主要发现

  • CAL-SHEQ在收敛速度与稳定性方面优于贝叶斯优化,实现了对交易量分位数与市场占有率等目标指标更平滑、更精确的校准。
  • CAL-SHEQ中的强化学习校准器实现了更高的平均奖励,并避免了负奖励——这与贝叶斯优化形成对比——通过确保智能体有足够时间适应新的超类型配置。
  • CAL-SHEQ中对超类型参数(如连接性、库存容忍度)的更新是平滑的,防止了贝叶斯优化基线中观察到的不稳定与奖励发散。
  • 该方法成功将商人智能体类型的分布校准至现实世界目标,如超类型1的第10百分位交易数量为8,实现了高精度与低方差。
  • 在多次实验中,收敛表现稳定且鲁棒,CAL-SHEQ在不同超类型配置(如2-2-2-3与3-4-5-5)下均表现出一致性能。
  • 双时间尺度框架确保在新调整前均衡已充分形成,从而实现可靠且一致的策略适应。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。