Skip to main content
QUICK REVIEW

[論文レビュー] Lyapunov stochastic stability and control of robust dynamic coalitional games with transferable utilities

Dario Bauso, Puduru Viswanadha Reddy|arXiv (Cornell University)|Jun 9, 2011
Game Theory and Voting Systems参考文献 19被引用数 3
ひとこと要約

本稿では、不確実性下での移転可能な利害対立ゲームにおけるロバストな動的割り当てルールを提案し、Lyapunov確率的安定性を用いて平均割り当てが長期平均ゲームのコアに収束することを保証する。この手法は、累積超過報酬(完全または部分的に観測可能)に基づくフィードバック制御を採用し、元の利害対立価値の確率分布を事前に知らない状況下でも、コアへの確実収束および超過量が事前に定義された錐内に収束することを保証する。

ABSTRACT

This paper considers a dynamic game with transferable utilities (TU), where the characteristic function is a continuous-time bounded mean ergodic process. A central planner interacts continuously over time with the players by choosing the instantaneous allocations subject to budget constraints. Before the game starts, the central planner knows the nature of the process (bounded mean ergodic), the bounded set from which the coalitions' values are sampled, and the long run average coalitions' values. On the other hand, he has no knowledge of the underlying probability function generating the coalitions' values. Our goal is to find allocation rules that use a measure of the extra reward that a coalition has received up to the current time by re-distributing the budget among the players. The objective is two-fold: i) guaranteeing convergence of the average allocations to the core (or a specific point in the core) of the average game, ii) driving the coalitions' excesses to an a priori given cone. The resulting allocation rules are robust as they guarantee the aforementioned convergence properties despite the uncertain and time-varying nature of the coaltions' values. We highlight three main contributions. First, we design an allocation rule based on full observation of the extra reward so that the average allocation approaches a specific point in the core of the average game, while the coalitions' excesses converge to an a priori given direction. Second, we design a new allocation rule based on partial observation on the extra reward so that the average allocation converges to the core of the average game, while the coalitions' excesses converge to an a priori given cone. And third, we establish connections to approachability theory and attainability theory.

研究の動機と目的

  • 利害対立価値が未知かつ時間変動する状況下でも、平均割り当てが長期平均ゲームのコアに収束することを保証する割り当てルールの設計。
  • 動的状況下での安定性と公平性を確保するため、利害対立の超過量を事前に指定された錐または方向に誘導する。
  • 元の確率分布を知らなくても、累積超過報酬の部分的または完全的観測に基づいて機能するロバストな制御則の開発。
  • Lyapunov確率的安定性と到達可能性・到達可能性理論の概念との間の理論的関係の確立。
  • 不確実性を有界平均エルゴード過程としてモデル化し、余剰蓄積に基づくフィードバック機構を用いることで、動的利害対立ゲームにおける安定性と公平性の確保。

提案手法

  • 中央計画者が、時刻 t までに利害対立が受領した累積超過報酬に基づいて割り当てを調整する動的フィードバック制御則を用いる。
  • 完全観測の場合、制御則は正確な累積超過ベクトルを用い、Lyapunov関数に基づく設計により、コア内の特定の点へシステムを誘導する。
  • 部分観測の場合、制御則は超過ベクトルの符号に基づく近似を用い、妥当性と安定性を保証するための利得パラメータ δ を導入する。
  • 制御則はLyapunov確率的安定性理論に基づき導出され、正規化された超過ベクトルのほとんど確実収束を用いて収束性が証明される。
  • 本手法は、特性関数の有界平均エルゴード性に依存しており、長期平均値が明示的に定義可能であり、事前に既知であることを保証する。
  • 予算制約と飽和関数を組み込み、割り当てが妥当な範囲内に保たれるようにし、名目割り当てと余剰を基準点として用いる。

実験結果

リサーチクエスチョン

  • RQ1利害対立価値が不確実かつ時間変動する状況下で、平均割り当てが平均ゲームのコアに収束するような割り当てルールはどのように設計可能か?
  • RQ2超過報酬の部分的または完全的観測下で、利害対立の超過量が事前に指定された錐または方向に収束することを保証する制御戦略は何か?
  • RQ3未知の確率分布を持つ動的利害対立ゲームにおいて、Lyapunov確率的安定性をどのように適用し、ほとんど確実収束を保証できるか?
  • RQ4累積超過報酬は、システムの安定化とコアへの割り当て誘導にどのように寄与するか?
  • RQ5提案された制御則は、確率的制御およびゲーム理論における到達可能性・到達可能性理論とどのように関係しているか?

主な発見

  • 超過報酬の完全観測に基づく制御則は、正規化された超過ベクトルがゼロにほとんど確実に収束することを保証し、これは平均割り当てがコア内の特定の点に収束することを意味する。
  • シミュレーション結果により、平均割り当てが長期的に名目割り当てベクトルに収束することが確認され、時間平均割り当てが名目値に収束する傾向が示された。
  • 部分観測制御則は、平均割り当てが確率的にコアに収束することを保証し、瞬時の割り当てが名目割り当ての近傍に位置することを示した。
  • シミュレーション結果は、コアリション {1,2} の正規化超過量が時間経過とともにゼロに収束することを示しており、完全および部分観測下での理論的収束を裏付けた。
  • δ = 1 の制御則は、最小および最大許容値で定義される範囲内に割り当てを保持することで妥当性を確保し、利得パラメータの保守的推定により検証された。
  • 理論的枠組みは、Lyapunov確率的安定性と到達可能性/到達可能性理論との間の正式な関係を確立し、動的利害対立ゲームの分析的基盤を強化した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。