Skip to main content
QUICK REVIEW

[论文解读] A distance for probability spaces, and long-term values in Markov Decision Processes and Repeated Games

Jérôme Renault, Xavier Venel|arXiv (Cornell University)|Feb 28, 2012
Markov Chains and Monte Carlo Methods参考文献 24被引用 12
一句话总结

本文在部分可观测马尔可夫决策过程(POMDPs)和重复博弈的信念空间上,引入了一种新的度量 $ d_* $,用于定义有限支撑的博雷尔概率测度空间。通过确保在 $ d_* $ 下转移为 1-利普希茨连续,作者建立了在任意评估函数下小怠慢情形的广义一致值的存在性与表征,将布罗克威尔的统一最优性推广至 POMDPs 和具有知情控制者的重复博弈。

ABSTRACT

Given a finite set $K$, we denote by $X=Δ(K)$ the set of probabilities on $K$ and by $Z=Δ_f(X)$ the set of Borel probabilities on $X$ with finite support. Studying a Markov Decision Process with partial information on $K$ naturally leads to a Markov Decision Process with full information on $X$. We introduce a new metric $d_*$ on $Z$ such that the transitions become 1-Lipschitz from $(X, \|.\|_1)$ to $(Z,d_*)$. In the first part of the article, we define and prove several properties of the metric $d_*$. Especially, $d_*$ satisfies a Kantorovich-Rubinstein type duality formula and can be characterized by using disintegrations. In the second part, we characterize the limit values in several classes of "compact non expansive" Markov Decision Processes. In particular we use the metric $d_*$ to characterize the limit value in Partial Observation MDP with finitely many states and in Repeated Games with an informed controller with finite sets of states and actions. Moreover in each case we can prove the existence of a generalized notion of uniform value where we consider not only the Cesàro mean when the number of stages is large enough but any evaluation function $θ\in Δ(\N^*)$ when the impatience $I(θ)=\sum_{t\geq 1} |θ_{t+1}-θ_t|$ is small enough.

研究动机与目标

  • 在 POMDPs 和重复博弈的信念空间上,定义有限支撑概率测度空间上的新度量 $ d_* $。
  • 证明在 $ d_* $ 下,转移变为 1-利普希茨连续,从而实现更强的收敛性与紧致性。
  • 在紧致非扩张的马尔可夫决策过程(特别是有限状态的 POMDPs 和具有知情控制者的重复博弈)中,表征极限值。
  • 证明在任意评估函数 $ \theta \in \Delta(\mathbb{N}^*) $ 且怠慢程度 $ I(\theta) \leq \alpha $ 的条件下,广义一致值存在。

提出的方法

  • 在 $ Z = \Delta_f(\Delta(K)) $ 上定义 $ d_* $,即信念空间 $ \Delta(K) $ 上的有限支撑博雷尔测度空间,利用分解与坎托罗维奇-鲁宾施特因对偶性。
  • 证明从 $ (X, \|\cdot\|_1) $ 到 $ (Z, d_*) $ 的转移动态为 1-利普希茨连续,确保稳定性和预紧性。
  • 利用度量 $ d_* $ 将 POMDPs 和重复博弈中的长期值分析转化为在 $ Z $ 上的完全信息 MDP。
  • 通过 $ d_* $-利普希茨转移的结构,利用块长上 Cesàro 平均的下确界表征广义一致值。
  • 在 $ Z \times [0,1] $ 上的辅助赌博屋模型中构造策略,这些策略可在原博弈中被模仿,以实现 $ \varepsilon $-最优性。
  • 应用表征式 $ v^*(\pi) = \inf_n \sup_m v_{m,n}(\pi) $,证明在小怠慢条件下,玩家 2 可保证获得 $ v^*(\pi) $。

实验结果

研究问题

  • RQ1能否在信念空间的有限支撑测度空间上定义一种新度量 $ d_* $,使得转移变为 1-利普希茨连续?
  • RQ2在任意评估函数下小怠慢情形中,POMDPs 和具有知情控制者的重复博弈中广义一致值是否存在?
  • RQ3此类博弈中的极限值能否通过块阶段上 Cesàro 平均的下确界来表征?
  • RQ4在单侧不完全信息的重复博弈中(例如 $ p \in [1/2, 1) $),是否可通过基于 $ d_* $ 的分析表征价值,尤其在无闭式解时?
  • RQ5能否在 $ Z \times [0,1] $ 上的辅助赌博屋模型中构造的策略,用于在原博弈中构造 $ \varepsilon $-最优策略?

主要发现

  • 度量 $ d_* $ 满足类似坎托罗维奇-鲁宾施特因的对偶性,且可通过测度的分解来表征。
  • 在 $ d_* $ 下,从 $ (X, \|\cdot\|_1) $ 到 $ (Z, d_*) $ 的转移映射为 1-利普希茨连续,确保了 $ Z $ 的稳定性和预紧性。
  • 广义一致值存在,并可表征为块长上 Cesàro 平均的下确界,即使在一般评估函数下小怠慢情形亦成立。
  • 对任意 $ \varepsilon > 0 $,存在 $ \alpha > 0 $,使得玩家 1 可在任意满足 $ I(\theta) \leq \alpha $ 的评估 $ \theta $ 下保证获得 $ v^*(\pi) - \varepsilon $。
  • 在相同条件下,玩家 2 也可通过独立于玩家 1 行动的分块最优策略保证获得 $ v^*(\pi) $。
  • 在单侧不完全信息重复博弈的示例中,当 $ p \in [1/2, 2/3) $ 时,价值为 $ v_p = \frac{p}{4p-1} $;在方程 $ 9p^3 - 12p^2 + 6p - 1 = 0 $ 的根 $ p^* $ 处,价值为 $ v_p = \frac{p^*}{1 - 3p^* + 6(p^*)^2} $,与理论一致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。