[论文解读] Lyapunov stochastic stability and control of robust dynamic coalitional games with transferable utilities
本文提出了一种在不确定性条件下针对具有可转移效用的合作博弈的鲁棒动态分配规则,利用李雅普诺夫随机稳定性确保平均分配收敛至长期平均博弈的核心。该方法基于累积超额收益(完全或部分观测)的反馈控制,即使在未知联盟价值底层概率分布的情况下,也能保证几乎必然收敛至核心,且超额值收敛至预设的锥形区域。
This paper considers a dynamic game with transferable utilities (TU), where the characteristic function is a continuous-time bounded mean ergodic process. A central planner interacts continuously over time with the players by choosing the instantaneous allocations subject to budget constraints. Before the game starts, the central planner knows the nature of the process (bounded mean ergodic), the bounded set from which the coalitions' values are sampled, and the long run average coalitions' values. On the other hand, he has no knowledge of the underlying probability function generating the coalitions' values. Our goal is to find allocation rules that use a measure of the extra reward that a coalition has received up to the current time by re-distributing the budget among the players. The objective is two-fold: i) guaranteeing convergence of the average allocations to the core (or a specific point in the core) of the average game, ii) driving the coalitions' excesses to an a priori given cone. The resulting allocation rules are robust as they guarantee the aforementioned convergence properties despite the uncertain and time-varying nature of the coaltions' values. We highlight three main contributions. First, we design an allocation rule based on full observation of the extra reward so that the average allocation approaches a specific point in the core of the average game, while the coalitions' excesses converge to an a priori given direction. Second, we design a new allocation rule based on partial observation on the extra reward so that the average allocation converges to the core of the average game, while the coalitions' excesses converge to an a priori given cone. And third, we establish connections to approachability theory and attainability theory.
研究动机与目标
- 设计分配规则,以确保在联盟价值未知且时变的情况下,平均分配仍能收敛至长期平均博弈的核心。
- 引导联盟的超额值收敛至预设的锥形区域或方向,确保在动态环境中的稳定性和公平性。
- 开发在部分或完全观测累积超额收益条件下仍能运行的鲁棒控制律,且无需依赖底层概率分布的知识。
- 建立李雅普诺夫随机稳定性与逼近性与可达性理论中概念之间的理论联系。
- 通过建模不确定性为有界均值遍历过程,并基于盈余累积的反馈机制,确保动态合作博弈中的稳定性和公平性。
提出的方法
- 中央规划者采用基于累积超额收益的动态反馈控制律,根据截至时间 t 联盟所获累积超额收益调整分配。
- 在完全观测下,控制律利用精确的累积超额向量,通过基于李雅普诺夫函数的设计,将系统引导至核心内的特定点。
- 在部分观测下,控制律使用超额向量的符号近似,并引入增益参数 δ 以确保可行性与稳定性。
- 控制律基于李雅普诺夫随机稳定性理论推导,收敛性通过归一化超额向量几乎必然收敛至零来证明。
- 该方法依赖于特征函数的有界均值遍历性,确保长期平均值有定义且可预先获知。
- 该框架结合预算约束与饱和函数,以保持分配在可行范围内,以名义分配和盈余作为参考点。
实验结果
研究问题
- RQ1当联盟价值不确定且时变时,如何设计分配规则,以确保平均分配收敛至平均博弈的核心?
- RQ2何种控制策略可保证在部分或完全观测超额收益的情况下,联盟的超额值收敛至预设的锥形区域或方向?
- RQ3如何应用李雅普诺夫随机稳定性,以确保在未知概率分布下,动态合作博弈中实现几乎必然收敛?
- RQ4累积超额收益在稳定系统并引导分配向核心靠拢的过程中起什么作用?
- RQ5所提出的控制律与随机控制与博弈论中逼近性与可达性理论之间有何关联?
主要发现
- 基于完全观测超额收益的控制律,可保证归一化超额向量几乎必然收敛至零,意味着平均分配收敛至核心内的特定点。
- 仿真结果证实,平均分配在长期收敛至名义分配向量,表现为时间平均分配收敛至名义值。
- 部分观测控制律确保平均分配以概率收敛至核心,且瞬时分配位于名义分配的邻域内。
- 仿真结果表明,联盟 {1,2} 的归一化超额值随时间收敛至零,验证了在完全与部分观测下理论收敛的有效性。
- 当 δ = 1 时,控制律通过保守估计增益参数,确保可行性,使分配保持在最小与最大允许值定义的边界内。
- 理论框架建立了李雅普诺夫随机稳定性与逼近性/可达性理论之间的正式联系,丰富了动态合作博弈的分析基础。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。