Skip to main content
QUICK REVIEW

[论文解读] Structure in the Value Function of Two-Player Zero-Sum Games of Incomplete Information

Auke Wiggers, Frans A. Oliehoek|arXiv (Cornell University)|Jun 22, 2016
Decision-Making and Behavioral Economics参考文献 16被引用 4
一句话总结

本文提出一种计划时充分统计量 σt,用于表示双人零和部分可观察马尔可夫决策过程(zs-POSGs)中的历史联合策略,通过逆向归纳实现理性决策。证明了价值函数 V*t(σt) 在代理1的类型分布上为凹函数,在代理2上为凸函数,为复杂不完全信息序列决策问题提供了高效求解方法的结构支撑。

ABSTRACT

Zero-sum stochastic games provide a rich model for competitive decision making. However, under general forms of state uncertainty as considered in the Partially Observable Stochastic Game (POSG), such decision making problems are still not very well understood. This paper makes a contribution to the theory of zero-sum POSGs by characterizing structure in their value function. In particular, we introduce a new formulation of the value function for zs-POSGs as a function of the "plan-time sufficient statistics" (roughly speaking the information distribution in the POSG), which has the potential to enable generalization over such information distributions. We further delineate this generalization capability by proving a structural result on the shape of value function: it exhibits concavity and convexity with respect to appropriately chosen marginals of the statistic space. This result is a key pre-cursor for developing solution methods that may be able to exploit such structure. Finally, we show how these results allow us to reduce a zs-POSG to a "centralized" model with shared observations, thereby transferring results for the latter, narrower class, to games with individual (private) observations.

研究动机与目标

  • 为解决双人零和POSGs中不完全信息下历史联合策略的表示问题。
  • 定义一个充分统计量σt,以捕捉未来阶段理性决策所需的所有相关历史信息。
  • 建立价值函数V*t(σt)在边际类型空间中具有凹/凸结构,从而实现高效优化。
  • 尽管充分统计量空间具有连续性和高维性,仍为zs-POSGs中的逆向归纳提供基础。
  • 正式验证σt足以用于价值计算,将协同Dec-POMDPs中的概念推广至零和设定。

提出的方法

  • 引入计划时充分统计量 σt(→θt) = Pr(→θt|b0, φt),表示在初始信念b0和历史联合策略φt下,对联合可观测历史(AOHs)的后验概率。
  • 定义σt的更新规则:σt+1(→θδt+1) ∝ Pr(o t+1 | →θδt, a t) · δt(a t | →θδt) · σt(→θδt)。
  • 通过充分统计量重述Q值函数Q*t(σt, →θδt, δt),最终阶段的Q值等于即时奖励R(→θδt, δt)。
  • 通过最大-最小优化推导理性决策规则:δt+1*1 = argmaxδ1t+1∈Δ1S minδ2t+1∈Δ2S Q*t+1(σt+1, ⟨δ1t+1, δ2t+1⟩),代理2同理。
  • 证明σt对价值计算是充分的:Q*t(σt, →θt, δt) = Q*t(φt, →θt, δt),确保从历史策略中不丢失信息。
  • 建立价值函数V*t(σt)在Δ(→Θ1t)上为凹函数,在Δ(→Θ2t)上为凸函数,利用最终阶段零和贝叶斯博弈族的结构。

实验结果

研究问题

  • RQ1能否定义一种计划时充分统计量σt,以捕捉双人零和POSG中所有历史联合策略的影响?
  • RQ2σt是否足以用于计算最优价值函数Q*t(σt, →θt, δt)而不丢失信息?
  • RQ3价值函数V*t(σt)是否在两名代理的边际类型分布中表现出凹/凸结构?
  • RQ4该结构特性能否被利用,以在连续且高维的统计量空间中实现zs-POSGs的高效逆向归纳?
  • RQ5最终阶段的价值函数与一组零和贝叶斯博弈有何关联?会涌现出何种结构特性?

主要发现

  • 计划时充分统计量σt被形式化证明对价值计算是充分的,即对所有t, →θt, δt,均有Q*t(σt, →θt, δt) = Q*t(φt, →θt, δt)。
  • 价值函数V*t(σt)在代理1的边际类型分布Δ(→Θ1t)上为凹函数,在代理2的Δ(→Θ2t)上为凸函数,从而支持高效优化。
  • zs-POSG的最终阶段(t = h−1)等价于一组零和贝叶斯博弈,其价值函数在各自代理的边际上继承了凹/凸结构。
  • 价值函数的结构允许分解为若干线性段,每段对应于给定对方策略时,某一代理的最佳响应策略。
  • 尽管σt具有连续性,凹/凸结构仍为可扩展算法提供了基础,尽管直接逆向归纳仍具挑战性。
  • 该方法将充分统计量的概念从协同Dec-POMDPs推广至零和设定,验证了其在不完全信息下理性决策中的适用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。