Skip to main content
QUICK REVIEW

[论文解读] A Predictive Theory of Games

David H. Wolpert|ArXiv.org|Dec 8, 2005
Game Theory and Applications参考文献 71被引用 3
一句话总结

本文提出了预测性博弈论(PGT),这是一种基于第一性原理的框架,用贝叶斯-概率方法替代均衡概念,以预测博弈中的联合策略。通过使用信息论先验和决策理论,PGT推导出联合策略的概率分布,表明纳什均衡不一定是预测性的,而定量响应均衡(QRE)则作为该分布众数的近似,其校正项通过最大熵方法推导得出。

ABSTRACT

Conventional noncooperative game theory hypothesizes that the joint strategy of a set of players in a game must satisfy an "equilibrium concept". All other joint strategies are considered impossible; the only issue is what equilibrium concept is "correct". This hypothesis violates the desiderata underlying probability theory. Indeed, probability theory renders moot the problem of what equilibrium concept is correct - every joint strategy can arise with non-zero probability. Rather than a first-principles derivation of an equilibrium concept, game theory requires a first-principles derivation of a distribution over joint (mixed) strategies. This paper shows how information theory can provide such a distribution over joint strategies. If a scientist external to the game wants to distill such a distribution to a point prediction, that prediction should be set by decision theory, using their (!) loss function. So the predicted joint strategy - the "equilibrium concept" - varies with the external scientist's loss function. It is shown here that in many games, having a probability distribution with support restricted to Nash equilibria - as stipulated by conventional game theory - is impossible. It is also show how to: i) Derive an information-theoretic quantification of a player's degree of rationality; ii) Derive bounded rationality as a cost of computation; iii) Elaborate the close formal relationship between game theory and statistical physics; iv) Use this relationship to extend game theory to allow stochastically varying numbers of players.

研究动机与目标

  • 挑战博弈论中将均衡概念视为唯一有效结果的传统依赖。
  • 解决基于均衡的模型与概率论基础公理之间的不一致性。
  • 基于信息论,发展一种从第一性原理推导联合策略分布的系统方法。
  • 表明纳什均衡并不总能通过理性极限逼近,且将支持集限制于纳什均衡在数学上常常不可行。
  • 通过统计力学类比,提供一种与具体模型无关的玩家行为理性程度量化方法。

提出的方法

  • 使用贝叶斯推断,基于博弈结构和先验信息,推导出联合混合策略的后验分布。
  • 在期望收益约束下应用最大熵(MaxEnt)原理,推导策略的概率分布。
  • 通过玻尔兹曼分布建模玩家行为,将理性程度与统计力学框架中的逆温度参数β关联。
  • 通过在分布众数附近展开,推导出对定量响应均衡(QRE)的校正项。
  • 利用决策理论,基于损失函数选择贝叶斯最优预测,从而无需依赖均衡概念。
  • 通过建模人口规模的涨落,将博弈论扩展至随机玩家数量,类比于统计物理中的系综。

实验结果

研究问题

  • RQ1为何对均衡概念的传统依赖与概率论基础公理存在不一致?
  • RQ2如何在博弈论中从第一性原理推导出联合策略的概率分布?
  • RQ3定量响应均衡(QRE)与所推导的策略分布众数之间存在何种正式关系?
  • RQ4是否每个纳什均衡都能作为理性程度不断提高的QRE策略序列的极限被逼近?若不能,原因是什么?
  • RQ5如何利用信息论度量,独立于特定模型,对玩家行为的理性程度进行量化?

主要发现

  • 在所推导分布下,非零概率的联合策略集合具有正测度,与均衡概念的零测度假设相矛盾。
  • 纳什均衡并不总能通过一系列理性程度不断提高的QRE策略逼近,且在许多博弈中,将分布的支持集限制于纳什均衡在数学上不可行。
  • 定量响应均衡(QRE)被证明是信息论分布下联合策略众数的近似,其校正项来自高阶展开。
  • 每个纳什均衡均可作为所有具有非零概率的联合策略序列的极限被逼近,尽管这些策略不一定是相关分布的众数。
  • 该框架通过逆温度参数β提供了与模型无关的理性程度量化,该参数自然地从期望收益约束下的最大熵推导中浮现。
  • 将框架扩展至随机玩家数量后,导致进化博弈论中复制子动态的修正,其本质是统计力学类比中人口规模的涨落。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。