Skip to main content
QUICK REVIEW

[论文解读] Asymptotic Convergence and Performance of Multi-Agent Q-Learning Dynamics

Aamal Hussain, Francesco Belardinelli|arXiv (Cornell University)|Jan 23, 2023
Experimental Behavioral Economics Studies被引用 5
一句话总结

本文建立了平滑Q-Learning动态中探索率的充分条件,确保在任意一般N人博弈中渐近收敛到唯一均衡。通过利用单调性和扰动分析,作者证明了加权势博弈和加权零和双矩阵博弈作为特例的收敛性,并进一步表明在特定条件下,非收敛动态在社会福利方面可优于均衡。

ABSTRACT

Achieving convergence of multiple learning agents in general $N$-player games is imperative for the development of safe and reliable machine learning (ML) algorithms and their application to autonomous systems. Yet it is known that, outside the bounds of simple two-player games, convergence cannot be taken for granted. To make progress in resolving this problem, we study the dynamics of smooth Q-Learning, a popular reinforcement learning algorithm which quantifies the tendency for learning agents to explore their state space or exploit their payoffs. We show a sufficient condition on the rate of exploration such that the Q-Learning dynamics is guaranteed to converge to a unique equilibrium in any game. We connect this result to games for which Q-Learning is known to converge with arbitrary exploration rates, including weighted Potential games and weighted zero sum polymatrix games. Finally, we examine the performance of the Q-Learning dynamic as measured by the Time Averaged Social Welfare, and comparing this with the Social Welfare achieved by the equilibrium. We provide a sufficient condition whereby the Q-Learning dynamic will outperform the equilibrium even if the dynamics do not converge.

研究动机与目标

  • 解决在一般N人博弈中确保多智能体强化学习收敛到唯一均衡的关键挑战。
  • 克服多智能体学习中常见的非平稳性和不稳定性问题,其中循环和混沌现象常出现。
  • 通过参数化探索率,为任意博弈中的平滑Q-Learning提供普遍收敛保证。
  • 通过时间平均社会福利分析性能,并与均衡表现进行比较,识别出非收敛动态更优的条件。
  • 基于单调性和扰动理论,将先前关于势博弈和零和博弈收敛性的结果统一到一个理论框架下。

提出的方法

  • 引入一种平滑Q-Learning动态模型,通过参数化探索率平衡探索与利用。
  • 应用扰动理论将原博弈转化为具有严格单调伪梯度的扰动博弈,确保收敛。
  • 利用加权单调性概念及凹势函数的性质,建立收敛条件。
  • 借助复制者动态(RD)及其在单调性下的收敛性质,通过扰动下的等价性推断Q-Learning的收敛性。
  • 定义并分析时间平均社会福利作为性能度量,用于比较学习动态与均衡结果。
  • 推导出在何种充分条件下,非收敛Q-Learning动态可产生高于均衡的社会福利,即使未实现收敛。
Asymptotic Convergence and Performance of Multi-Agent Q-Learning Dynamics

实验结果

研究问题

  • RQ1在何种探索率条件下,平滑Q-Learning在任意一般N人博弈中收敛到唯一均衡?
  • RQ2特定博弈类别的收敛结果——如加权势博弈和加权零和双矩阵博弈——能否在单一理论框架下统一?
  • RQ3是否存在非收敛Q-Learning动态优于均衡结果的社会福利的情境?
  • RQ4以时间平均社会福利衡量的Q-Learning动态性能,与均衡处的社会福利相比如何?
  • RQ5能否通过参数调节保持或诱导博弈伪梯度的单调性,以确保收敛?

主要发现

  • 探索率的充分条件——依赖于博弈规模和玩家数量——可确保平滑Q-Learning动态在任意N人博弈中收敛到唯一均衡。
  • 具有凹势函数的加权势博弈和加权零和双矩阵博弈被证明是主收敛结果的特例,其收敛性通过单调性和扰动分析得到证明。
  • 当势函数为凹函数时,扰动博弈 $Γ^H$ 具有严格单调的伪梯度,从而通过复制者动态的既定结果确保Q-Learning收敛。
  • 对于加权零和双矩阵博弈,由于零和性质和对称收益结构,博弈被证明具有加权单调性,从而收敛到唯一的量化响应均衡(QRE)。
  • 当动态不收敛时,存在一个充分条件,使得Q-Learning的时间平均社会福利超过均衡水平,表明避免完全收敛可能具有优势。
  • 实验表明,增加探索可能降低收益表现,提示在某些博弈中,最小化或完全不进行探索可能比将智能体推向均衡产生更好的社会福利。
Asymptotic Convergence and Performance of Multi-Agent Q-Learning Dynamics

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。