Skip to main content
QUICK REVIEW

[论文解读] Learning in concave games with imperfect information.

Panayotis Mertikopoulos|arXiv (Cornell University)|Aug 25, 2016
Experimental Behavioral Economics Studies被引用 6
一句话总结

本文研究了在不完全信息下凹型N人博弈中的学习动态,其中玩家通过在有噪声的收益估计上进行梯度上升来更新动作,并将结果投影回可行集。在收益梯度噪声有界的条件下,证明了收敛到纳什均衡;建立了变分稳定性作为高概率下局部收敛的充分条件;并表明在满射镜像映射下,即使存在不确定性,也能实现严格均衡的有限时间收敛。

ABSTRACT

This paper examines the convergence properties of a class of learning schemes for concave N-person games - that is, games with convex action spaces and individually concave payoff functions. Specifically, we focus on a family of learning methods where players adjust their actions by taking small steps along their individual payoff gradients and then the output back to their feasible action spaces. Assuming players only have access to gradient information that is accurate up to a zero-mean error with bounded variance, we show that when the process converges, its limit is a Nash equilibrium. We also introduce an equilibrium stability notion which we call variational stability (VS), and we show that stable equilibria are locally attracting with high probability whereas globally stable states are globally attracting with probability 1. Additionally, in finite games, we find that dominated strategies become extinct, strict equilibria are locally attracting with high probability, and the long-term average of the process converges to equilibrium in 2-player zero-sum games. Finally, we examine the scheme's convergence speed and we show that if the game admits a strict equilibrium and the players' mirror maps are surjective, then, with high probability, the process converges to equilibrium in a finite number of steps, no matter the level of uncertainty.

研究动机与目标

  • 分析在不完全信息下基于梯度的学习在凹型N人博弈中的收敛特性。
  • 建立在收益梯度受零均值、有界方差噪声影响下,学习动态收敛到纳什均衡的条件。
  • 引入并分析一种新稳定性概念——变分稳定性(VS),及其对收敛性的影响。
  • 研究有限博弈中的长期行为,包括被支配策略的消失以及两人零和博弈中收敛到均衡的情况。
  • 刻画收敛速度,特别是在严格均衡和满射镜像映射条件下的表现。

提出的方法

  • 玩家通过在具有有界方差、零均值噪声的估计收益梯度上进行小幅步长更新来调整动作。
  • 每次更新后,利用镜像映射将动作投影回可行动作空间。
  • 分析依赖于随机逼近理论,以在不确定性下建模学习动态。
  • 引入一种新的稳定性概念——变分稳定性(VS),用于刻画局部吸引的均衡。
  • 使用类似李雅普诺夫的论证方法以及对学习轨迹的概率界来研究收敛性。
  • 将该框架应用于有限博弈和两人零和博弈,推导出长期收敛结果。

实验结果

研究问题

  • RQ1当收益梯度受零均值、有界方差噪声污染时,基于梯度的学习在凹型博弈中在何种条件下收敛到纳什均衡?
  • RQ2所引入的变分稳定性(VS)概念与学习动态中均衡的局部吸引性之间有何关系?
  • RQ3在此学习方案下,有限博弈中的被支配策略和严格均衡会发生什么变化?
  • RQ4学习过程是否在两人零和博弈中收敛到均衡?其平均动作路径的长期行为如何?
  • RQ5在满射镜像映射和噪声梯度条件下,能否保证有限时间收敛到严格均衡?

主要发现

  • 当学习过程收敛时,在收益梯度估计的噪声有界条件下,其极限为纳什均衡。
  • 变分稳定均衡在高概率下是局部吸引的,而全局稳定均衡则以概率1全局吸引。
  • 在有限博弈中,被支配策略在学习动态下会趋于消失。
  • 严格均衡在高概率下是局部吸引的,且在两人零和博弈中,该过程的长期平均收敛到均衡。
  • 若存在严格均衡且玩家的镜像映射为满射,则无论噪声水平如何,该过程以高概率在有限步内收敛到均衡。
  • 收敛速度被刻画为:在特定结构条件下,即使存在显著不确定性,也可能实现有限时间收敛。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。