Skip to main content
QUICK REVIEW

[论文解读] Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous Time

Wei-Chen Wang, Jiequn Han|arXiv (Cornell University)|Aug 16, 2020
Reinforcement Learning in Robotics参考文献 51被引用 6
一句话总结

本文建立了连续时间线性二次均场控制与博弈问题中策略梯度方法的全局收敛性,证明在温和条件下具有线性收敛速率。通过为Riccati系统与线性系统建立新颖的稳定性与灵敏度界,将离散时间策略梯度分析扩展至连续时间随机动力系统。

ABSTRACT

Reinforcement learning is a powerful tool to learn the optimal policy of possibly multiple agents by interacting with the environment. As the number of agents grow to be very large, the system can be approximated by a mean-field problem. Therefore, it has motivated new research directions for mean-field control (MFC) and mean-field game (MFG). In this paper, we study the policy gradient method for the linear-quadratic mean-field control and game, where we assume each agent has identical linear state transitions and quadratic cost functions. While most of the recent works on policy gradient for MFC and MFG are based on discrete-time models, we focus on the continuous-time models where some analyzing techniques can be interesting to the readers. For both MFC and MFG, we provide policy gradient update and show that it converges to the optimal solution at a linear rate, which is verified by a synthetic simulation. For MFG, we also provide sufficient conditions for the existence and uniqueness of the Nash equilibrium.

研究动机与目标

  • 填补连续时间均场控制与博弈(MFC/MFG)设置下策略梯度方法理论收敛分析的空白。
  • 将现有离散时间策略梯度收敛结果扩展至具有线性二次结构的连续时间随机动力系统。
  • 为均场控制(MFC)与均场博弈(MFG)公式提供可证明的全局线性收敛性。
  • 建立MFG设置下纳什均衡存在与唯一的充分条件。
  • 提出一种两步算法框架用于MFG,通过交替求解最优响应与更新均场状态,确保收敛。

提出的方法

  • 通过参数重参数化,将均场控制(MFC)问题重新表述为标准的线性二次调节器(LQR)问题。
  • 引入漂移LQR问题作为中间步骤,以分析更复杂的MFG设置。
  • 设计一种策略梯度算法,交替求解最优响应问题(漂移LQR)与更新均场状态。
  • 利用系统参数的利普希茨连续性与逆矩阵灵敏度,推导Riccati系统与线性系统的稳定性界。
  • 通过递归误差分解方法,利用谱范数与条件数,界定向量迭代与最优策略之间的距离。
  • 基于问题特定常数设定自适应步长(ε_s),确保在最优解的ε-邻域内收敛。

实验结果

研究问题

  • RQ1策略梯度方法能否在连续时间线性二次均场控制问题中实现全局收敛?
  • RQ2策略梯度在连续时间MFC与MFG设置下的收敛速率如何?
  • RQ3如何将离散时间策略梯度的理论分析扩展至连续时间随机动力系统?
  • RQ4在连续时间均场博弈中,纳什均衡在何种条件下存在且唯一?
  • RQ5何种算法结构可确保在耦合均场动力系统的MFG设置中实现线性收敛?

主要发现

  • 连续时间线性二次均场控制的策略梯度方法以线性速率全局收敛至最优解。
  • 对于均场博弈,所提出的两步算法在存在与唯一性充分条件下,实现向纳什均衡的线性收敛。
  • 收敛速率由几何衰减因子L₀ < 1量化,确保在S次迭代后,误差被限制在ε以内。
  • 算法在S = O(log(1/ε))次迭代后达到策略参数(K与b)的ε-精度,其中S依赖于初始误差与L₀。
  • 参数误差界依赖于系统矩阵(R, D, Dᵀ)的谱性质,显式依赖于σ_min(R)与σ_min(DDᵀ)。
  • 收敛对扰动具有鲁棒性,误差界包含C_b(μ_s)与C_K(μ_s),其随系统与策略灵敏度而缩放。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。