Skip to main content
QUICK REVIEW

[论文解读] A tutorial on Zero-sum Stochastic Games

Jérôme Renault|arXiv (Cornell University)|May 16, 2019
Economic theories and models参考文献 20被引用 5
一句话总结

本教程对零和随机博弈提供了一个全面的数学导引,这是马尔可夫决策过程在竞争性双人动态设定下的扩展。它建立了n阶段博弈和λ-折扣博弈中值的存在性与收敛性的基础结果,证明了统一值的存在性,并展示了在连续收益和转移结构下,n阶段博弈与λ-折扣博弈的渐近行为之间的等价性。

ABSTRACT

Zero-sum stochastic games generalize the notion of Markov Decision Processes (i.e. controlled Markov chains, or stochastic dynamic programming) to the 2-player competitive case : two players jointly control the evolution of a state variable, and have opposite interests. These notes constitute a short mathematical introduction to the theory of such games. Section 1 presents the basic model with finitely many states and actions. We give proofs of the standard results concerning : the existence and formulas for the values of the n-stage games, of the $λ$-discounted games, the convergence of these values when $λ$ goes to 0 (algebraic approach) and when n goes to +$\infty$, an important example called 'The Big Match' and the existence of the uniform value. Section 2 presents a short and subjective selection of related and more recent results : 1-player games (MDP) and the compact non expansive case, a simple compact continuous stochastic game with no asymptotic value, and the general equivalence between the uniform convergence of (v n) n and (v $λ$) $λ$. More references on the topic can be found for instance in the books by Mertens-Sorin

研究动机与目标

  • 为零和随机博弈提供严谨的数学基础,将马尔可夫决策过程推广至竞争性双人设定。
  • 在有限状态空间和动作空间下,建立n阶段博弈和λ-折扣博弈的值的存在性与公式。
  • 分析当n → ∞和λ → 0时值的收敛性,证明在一般情况下统一值的存在性。
  • 探讨在连续随机博弈中,n阶段博弈与λ-折扣博弈值序列的渐近行为之间的等价性。
  • 通过《大比赛》和吸收性博弈等关键例子,阐明理论概念并揭示开放问题。

提出的方法

  • 使用有限的状态集K、动作集I和J、收益函数g: K×I×J → ℝ,以及转移概率q: K×I×J → Δ(K)来形式化零和随机博弈的基本模型。
  • 定义策略,包括行为策略、马尔可夫策略和平稳策略,并利用乘积σ-代数在无限对弈路径上定义概率测度。
  • 应用夏普利方程,通过在混合动作上的上确界-下确界优化,递归地刻画值函数v_n(k)和v_λ(k)。
  • 采用代数与渐近分析方法,研究当n → ∞和λ → 0时,v_n(k)和v_λ(k)的极限行为。
  • 利用紧致性与连续性论证,证明在连续随机博弈中,(v_n)与(v_λ)的统一收敛性等价。
  • 通过分析《大比赛》和单人博弈等具体例子,说明非平凡的收敛行为以及在某些情况下渐近值的缺失。

实验结果

研究问题

  • RQ1在何种条件下,n阶段博弈与λ-折扣博弈的值在n → ∞和λ → 0时收敛?
  • RQ2在连续收益与转移结构下,零和随机博弈中统一值的存在性是否可由v_n与v_λ的收敛性推出?
  • RQ3在一般随机博弈中,即使值不收敛,v_n与v_λ的渐近行为是否仍可等价?
  • RQ4在某些连续随机博弈中,渐近值的缺失具有何种含义?
  • RQ5吸收态的性质以及《大比赛》等特定博弈结构如何影响长期值的收敛性?

主要发现

  • 对于所有n和λ ∈ (0,1),n阶段博弈的值v_n(k)与λ-折扣博弈的值v_λ(k)均存在,且满足夏普利方程。
  • 当n → ∞时,值v_n(k)收敛于统一值,该值在给定假设下存在且定义良好。
  • 当λ → 0时,值v_λ(k)收敛于与n → ∞时v_n(k)相同的极限,从而证明了两种渐近行为的等价性。
  • 在紧致度量状态空间与动作空间,以及连续收益与转移函数的连续情况下,(v_n)与(v_λ)的统一收敛性等价。
  • 对于《大比赛》,当λ → 0时,值v_λ(k)不收敛,表明在某些情况下渐近值可能不存在。
  • v_λ(k)的极限上界为1/2;在特定λ_n序列下,y_λ与x_λ的极限下确界与上确界分别为4/9与1/2,从而证明了该情况下v_λ(k)的非收敛性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。