Skip to main content
QUICK REVIEW

[论文解读] Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms

Kaiqing Zhang, Zhuoran Yang|arXiv (Cornell University)|Nov 24, 2019
Innovation Diffusion and Forecasting被引用 124
一句话总结

对MARL的理论综述,聚焦两个框架(马尔可夫/随机博弈与扩展式博弈),分析收敛性、复杂性,以及未来研究的新视角。

ABSTRACT

Recent years have witnessed significant advances in reinforcement learning (RL), which has registered great success in solving various sequential decision-making problems in machine learning. Most of the successful RL applications, e.g., the games of Go and Poker, robotics, and autonomous driving, involve the participation of more than one single agent, which naturally fall into the realm of multi-agent RL (MARL), a domain with a relatively long history, and has recently re-emerged due to advances in single-agent RL techniques. Though empirically successful, theoretical foundations for MARL are relatively lacking in the literature. In this chapter, we provide a selective overview of MARL, with focus on algorithms backed by theoretical analysis. More specifically, we review the theoretical results of MARL algorithms mainly within two representative frameworks, Markov/stochastic games and extensive-form games, in accordance with the types of tasks they address, i.e., fully cooperative, fully competitive, and a mix of the two. We also introduce several significant but challenging applications of these algorithms. Orthogonal to the existing reviews on MARL, we highlight several new angles and taxonomies of MARL theory, including learning in extensive-form games, decentralized MARL with networked agents, MARL in the mean-field regime, (non-)convergence of policy-based methods for learning in games, etc. Some of the new angles extrapolate from our own research endeavors and interests. Our overall goal with this chapter is, beyond providing an assessment of the current state of the field on the mark, to identify fruitful future research directions on theoretical studies of MARL. We expect this chapter to serve as continuing stimulus for researchers interested in working on this exciting while challenging topic.

研究动机与目标

  • 澄清代表性框架(马尔可夫/随机博弈与扩展式博弈)下 MARL 的理论基础。
  • 在完全合作、完全竞争和混合设置下,组织具有收敛性与复杂性分析的 MARL 算法。
  • 突出 MARL 理论中的新角度与分类法,以指导未来研究与应用。

提出的方法

  • 在马尔可夫/随机与扩展式博弈框架内对具有理论保 证的 MARL 算法进行综述和综合。
  • 讨论诸如非平稳性、联合行动空间和信息结构等挑战,并将其与均衡概念联系起来。
  • 引入并比较设定(合作、竞争、混合)及其对学习动力学与收敛性的影响。
  • 突出扩展,如去中心化 MARL、平均场 MARL,以及在扩展式博弈中学习。

实验结果

研究问题

  • RQ1在马尔可夫/随机与扩展式博弈框架下,哪些 MARL 算法具有理论收敛性和复杂性保证?
  • RQ2合作、竞争和混合设定如何影响 MARL 的学习动力学与均衡概念?
  • RQ3MARL 理论中有哪些新兴角度和分类法可以指导未来的理论工作?

主要发现

  • 本章对 MARL 理论与算法进行了有选择性的概述,重点在框架与收敛性分析。
  • 它讨论了 MARL 常见的挑战,如非平稳性、组合性的联合行动空间和信息结构,并将它们与均衡概念联系起来。
  • 它将讨论扩展到扩展式博弈、去中心化 MARL、平均场 MARL,以及策略基方法在零和博弈中的(非)收敛性。
  • 它将纳什均衡和 ε-纳什均衡视为马尔可夫与扩展式博弈 MARL 设置中的核心解概念。
  • 它区分合作、竞争和混合设定,并展示它们如何映射到 MGs 和扩展式博弈,指导算法设计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。