[论文解读] A Hybrid Optimal Control Approach to LQG Mean Field Games with Switching and Stopping Strategies.
本文提出了一种混合最优控制框架,结合了平均场博弈(MFG)与线性二次高斯(LQG)理论,以在具有一个主要代理和两个次要群体的非合作随机博弈中推导出 $ 5c $-纳什均衡。该方法通过离散状态对切换动态和停止时间进行建模,得出所有代理在二次成本准则下的最优切换与停止策略以及最佳响应控制。
A novel framework is presented that combines Mean Field Game (MFG) theory and Hybrid Optimal Control (HOC) theory to obtain a unique $\epsilon$-Nash equilibrium for a non-cooperative game with stopping times. We consider the case where there exists one major agent with a significant influence on the system together with a large number of minor agents constituting two subpopulations, each with individually asymptotically negligible effect on the whole system. Each agent has stochastic linear dynamics with quadratic costs, and the agents are coupled in their dynamics by the average state of minor agents (i.e. the empirical mean field). The hybrid feature enters via the indexing by discrete states: (i) the switching of the major agent between alternative dynamics or (ii) the termination of the agents' trajectories in one or both of the subpopulations of minor agents. Optimal switchings and stopping time strategies together with best response control actions for, respectively, the major agent and all minor agents are established with respect to their individual cost criteria by an application of LQG HOC theory.
研究动机与目标
- 解决涉及一个主要代理和大量对系统影响渐近可忽略的次要代理的非合作随机博弈问题。
- 通过离散状态索引对代理经历切换动态或终止(停止时间)的系统进行建模。
- 在二次成本函数下,为主要代理和次要代理建立最优控制策略——包括切换与停止策略。
- 通过结合 LQG 与混合最优控制理论,推导出唯一的 $ 5c$-纳什均衡。
- 通过将代理动态与次要代理的经验平均场耦合,确保均衡的鲁棒性。
提出的方法
- 构建一种混合最优控制(HOC)结构,其中离散状态表示在不同动态之间的切换或代理轨迹的终止。
- 应用 LQG 理论,在线性随机动态和二次成本函数下求解最优控制问题。
- 通过经验平均场对代理之间的耦合进行建模,表示次要代理的平均状态。
- 利用动态规划和混合设置下的哈密顿-雅可比-贝尔曼方程,推导所有代理的最佳响应控制动作。
- 通过 MFG 与 HOC 框架的相互作用,建立 $ 5c$-纳什均衡的存在性。
- 采用两子群体结构:次要代理被划分为两组,每组具有不同的动态和停止行为。
实验结果
研究问题
- RQ1如何将混合最优控制框架扩展以在具有切换与停止策略的系统中纳入平均场博弈动态?
- RQ2在具有一个主要代理和两个次要群体的非合作博弈中,何种条件可确保唯一 $ 5c$-纳什均衡的存在?
- RQ3在二次成本准则下,切换动态与停止时间如何影响主要代理和次要代理的最优控制策略?
- RQ4次要代理的经验平均场在混合 LQG 设置下如何影响均衡策略?
- RQ5LQG 与 HOC 理论的结合能否为大规模群体随机博弈中的切换与停止行为提供可处理且最优的控制策略?
主要发现
- 该框架成功为包含一个主要代理和两个次要群体的非合作博弈建立了唯一的 $ 5c$-纳什均衡。
- 通过整合 LQG 与 HOC 理论,分别推导出主要代理和次要代理的最优切换与停止时间策略。
- 通过次要代理的经验平均状态实现的平均场耦合,确保了个体影响的渐近可忽略性,同时保持了系统范围内的相互作用。
- 在混合控制设置下,利用动态规划和哈密顿-雅可比-贝尔曼方程,明确刻画了所有代理的最佳响应控制动作。
- 混合结构通过引入用于切换和终止的离散状态,能够对大规模随机博弈中复杂代理行为进行建模。
- 该解决方案框架具有可扩展性,适用于表现出切换动态和有限时域停止行为的系统。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。