Skip to main content
QUICK REVIEW

[论文解读] Existence of the uniform value in repeated games with a more informed controller

Fabien Gensbittel, Miquel Oliu‐Barton|arXiv (Cornell University)|Jan 9, 2013
Game Theory and Applications参考文献 8被引用 8
一句话总结

该论文证明了在零和重复博弈中,当一名玩家(玩家1)对状态更为了解并控制另一名玩家信念演化时,一致值的存在性。通过构建一个基于二阶信念(玩家2对玩家1信念的信念)的辅助随机博弈,并利用Renault关于动态规划的存在性结果,作者证明了玩家1可将该辅助博弈中的最优策略实施于原博弈中,从而在包括不完全监控在内的广泛条件下确保一致值的存在。

ABSTRACT

We prove that in a general zero-sum repeated game where the first player is more informed than the second player and controls the evolution of information on the state, the uniform value exists. This result extends previous results on Markov decision processes with partial observation (Rosenberg, Solan, Vieille 2002), and repeated games with an informed controller (Renault 2012). Our formal definition of a more informed player is more general than the inclusion of signals, allowing therefore for imperfect monitoring of actions. We construct an auxiliary stochastic game whose state space is the set of second order beliefs of player 2 (beliefs about beliefs of player 1 on the true state variable of the initial game) with perfect monitoring and we prove it has a value by using a result of Renault 2012. A key element in this work is to prove that player 1 can use strategies of the auxiliary game in the initial game in our general framework, which allows to deduce that the value of the auxiliary game is also the value of our initial repeated game by using classical arguments.

研究动机与目标

  • 证明在零和重复博弈中,当玩家1更为了解且控制玩家2信念演化时,一致值的存在性。
  • 推广先前关于部分可观测马尔可夫决策过程和知情控制者重复博弈的结果。
  • 形式化一个更广泛的“更为了解”的概念,以允许对行动的不完全监控。
  • 证明可从二阶信念的辅助博弈中转移策略至原博弈。
  • 表明双方均可保证相同的值,从而证明一致值的存在。

提出的方法

  • 构建一个辅助随机博弈,其状态空间为二阶信念集合——即玩家2对玩家1关于真实状态信念的信念。
  • 利用Renault关于动态规划的结果,证明该辅助博弈具有一致值。
  • 通过信念一致性与策略转移,证明玩家1可将在辅助博弈中获得的最优策略实施于原博弈中。
  • 通过阶段收益的递归论证,证明玩家2可通过分块策略保证相同的值。
  • 运用经典博弈论论证,将辅助博弈的值与原博弈的一致值等同起来。
  • 应用假设(A1)–(A3),以确保玩家1可在不知晓玩家2策略的情况下计算信念,并确保信念演化的一致性。

实验结果

研究问题

  • RQ1在一名玩家更为了解并控制另一名玩家信念演化的过程中,重复博弈中的一致值是否存在?
  • RQ2是否可从基于二阶信念的辅助博弈中推导出不完全信息重复博弈的值?
  • RQ3在何种条件下,可将具有完全监控的辅助博弈中的策略用于具有不完全监控的博弈中?
  • RQ4是否可能将一致值存在性的结果扩展至不完全监控行动的模型之外?
  • RQ5玩家2是否可通过分块策略在这些博弈中保证与一致值相同的值?

主要发现

  • 在零和重复博弈中,当玩家1更为了解并控制玩家2信念演化时,即使在不完全监控下,一致值也存在。
  • 原博弈的值等于一个辅助随机博弈的值,其状态空间为玩家2的二阶信念集合。
  • 玩家1可将在辅助博弈中获得的最优策略实施于原博弈中,从而确保该值可实现。
  • 玩家2可通过分块策略,利用仅依赖于当前阶段和信念状态的递归策略,保证一致值。
  • 一致值的存在性在弱于先前猜想的假设下依然成立,包括行动不完全监控的情形。
  • 该结果推广了先前关于部分可观测马尔可夫决策过程和完全知情控制者重复博弈的研究成果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。