[论文解读] Learning Parametric Closed-Loop Policies for Markov Potential Games
本文提出了一种新颖的方法,用于在具有连续状态、连续动作和耦合约束的马尔可夫势博弈(MPG)中学习参数化闭环策略。通过将问题重新表述为单个最优控制问题(OCP),即使在复杂且非凸的奖励函数以及神经网络策略下,也能利用深度强化学习高效逼近闭环纳什均衡,克服了以往开环或变分方法的局限性。
Multiagent systems where agents interact among themselves and with a stochastic environment can be formalized as stochastic games. We study a subclass named Markov potential games (MPGs) that appear often in economic and engineering applications when the agents share a common resource. We consider MPGs with continuous state-action variables, coupled constraints and nonconvex rewards. Previous analysis followed a variational approach that is only valid for very simple cases (convex rewards, invertible dynamics, and no coupled constraints); or considered deterministic dynamics and provided open-loop (OL) analysis, studying strategies that consist in predefined action sequences, which are not optimal for stochastic environments. We present a closed-loop (CL) analysis for MPGs and consider parametric policies that depend on the current state. We provide easily verifiable, sufficient and necessary conditions for a stochastic game to be an MPG, even for complex parametric functions (e.g., deep neural networks); and show that a closed-loop Nash equilibrium (NE) can be found (or at least approximated) by solving a related optimal control problem (OCP). This is useful since solving an OCP--which is a single-objective problem--is usually much simpler than solving the original set of coupled OCPs that form the game--which is a multiobjective control problem. This is a considerable improvement over the previously standard approach for the CL analysis of MPGs, which gives no approximate solution if no NE belongs to the chosen parametric family, and which is practical only for simple parametric forms. We illustrate the theoretical contributions with an example by applying our approach to a noncooperative communications engineering game. We then solve the game with a deep reinforcement learning algorithm that learns policies that closely approximates an exact variational NE of the game.
研究动机与目标
- 解决在具有耦合约束的连续状态、连续动作马尔可夫势博弈(MPG)中,计算或近似闭环纳什均衡(CL-NE)的实用方法缺乏的问题。
- 将先前局限于凸奖励或简单参数形式的变分分析与开环分析扩展至一般性、非凸且复杂的参数化策略(如深度神经网络)。
- 为即使在非凸性和复杂策略参数化下,也提供足够且必要的条件,使一个随机博弈成为MPG。
- 证明求解单个最优控制问题(OCP)可得到等价于CL-NE的解,显著简化了原始的耦合多目标OCP系统。
- 在真实世界的通信工程博弈中验证该方法,展示通过深度强化学习实现接近最优均衡的收敛性。
提出的方法
- 将MPG表述为参数化闭环博弈,其中每个智能体的策略是当前状态的函数,由向量 $ w_k $ 参数化。
- 引入参数化闭环纳什均衡(PCL-NE)的概念,定义为策略向量 $ w^* $,其中任一智能体单方面偏离均无法获益。
- 利用一阶最优性条件,推导出使博弈成为MPG的充分必要条件,即使在非凸奖励和复杂参数函数下也成立。
- 通过利用势博弈结构,将博弈的耦合多目标OCP系统重表述为单个OCP,从而可借助标准最优控制求解器高效求解。
- 使用深度强化学习训练策略以近似PCL-NE,策略网络架构支持高表达能力(如深度神经网络)。
- 将该方法应用于非合作通信工程博弈,通过实证评估验证其有效性。
实验结果
研究问题
- RQ1我们能否为具有连续变量和耦合约束的随机博弈,推导出可验证的、充分且必要的条件,使其成为马尔可夫势博弈(MPG),即使在非凸奖励和复杂策略下?
- RQ2我们能否使用参数化策略近似MPG中的闭环纳什均衡(CL-NE),特别是在精确NE不在所选策略族中的情况下?
- RQ3将MPG的耦合多目标OCP系统重表述为单个OCP,是否能实现比求解原始耦合系统更高效、更可扩展的均衡计算?
- RQ4深度强化学习能否有效学习策略,使其在复杂的真实世界MPG应用中紧密逼近真实变分NE?
- RQ5在随机、受限环境中,与以往的开环或变分方法相比,该方法在性能和鲁棒性方面表现如何?
主要发现
- 本文建立了易于验证的、充分且必要的条件,用于判断一个随机博弈是否为MPG,即使在非凸奖励和复杂策略(如深度神经网络)下也成立。
- 通过求解单个最优控制问题(OCP),可找到或近似闭环纳什均衡(CL-NE),其计算复杂度远低于求解原始的耦合OCP系统。
- 所提方法实现了在具有连续状态、动作和耦合约束的MPG中对CL-NE的实际计算——此前标准方法难以处理此类问题。
- 在非合作通信工程博弈中,深度强化学习算法成功学习到的策略紧密逼近了精确的变分NE,证明了该方法的实证有效性。
- 该方法超越了简单参数形式的限制,允许使用高容量策略(如深度网络)实现对真实NE的任意接近逼近,前提是策略族具备足够的表达能力。
- 该方法克服了以往闭环分析的关键局限:当精确NE不在策略类中时,无法提供近似解;而本方法即使在复杂、非可逆动力学下也具有实用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。