Skip to main content
QUICK REVIEW

[论文解读] Incentive and stability in the Rock-Paper-Scissors game: an experimental investigation

Zhijian Wang, Bin Xu|arXiv (Cornell University)|Jul 4, 2014
Evolutionary Game Theory and Cooperation参考文献 56被引用 5
一句话总结

本实验研究探讨了在广义石头剪刀布游戏中,奖励水平(获胜收益 $a$)如何影响个体与集体策略动态。通过84组共720轮对弈,研究发现随着 $a$ 增大,最优响应行为增加,而赢则保持输则改变(WSLS)行为减少,且在 $a=2$ 附近出现明显的相变,揭示了人类学习策略随收益结构系统性转变。

ABSTRACT

In a two-person Rock-Paper-Scissors (RPS) game, if we set a loss worth nothing and a tie worth 1, and the payoff of winning (the incentive a) as a variable, this game is called as generalized RPS game. The generalized RPS game is a representative mathematical model to illustrate the game dynamics, appearing widely in textbook. However, how actual motions in these games depend on the incentive has never been reported quantitatively. Using the data from 7 games with different incentives, including 84 groups of 6 subjects playing the game in 300-round, with random-pair tournaments and local information recorded, we find that, both on social and individual level, the actual motions are changing continuously with the incentive. More expressively, some representative findings are, (1) in social collective strategy transit views, the forward transition vector field is more and more centripetal as the stability of the system increasing; (2) In the individual behavior of strategy transit view, there exists a phase transformation as the stability of the systems increasing, and the phase transformation point being near the standard RPS; (3) Conditional response behaviors are structurally changing accompanied by the controlled incentive. As a whole, the best response behavior increases and the win-stay lose-shift (WSLS) behavior declines with the incentive. Further, the outcome of win, tie, and lose influence the best response behavior and WSLS behavior. Both as the best response behavior, the win-stay behavior declines with the incentive while the lose-left-shift behavior increase with the incentive. And both as the WSLS behavior, the lose-left-shift behavior increase with the incentive, but the lose-right-shift behaviors declines with the incentive. We hope to learn which one in tens of learning models can interpret the empirical observation above.

研究动机与目标

  • 通过实证方法检验获胜收益 $a$ 的变化如何影响广义石头剪刀布游戏中策略动态。
  • 探究个体学习行为(尤其是最优响应与WSLS)是否随激励水平系统性变化。
  • 识别条件响应行为的结构性变化作为 $a$ 的函数,特别是在理论稳定性阈值 $a=2$ 附近。
  • 检验既有的学习模型是否能解释不同激励水平下观察到的行为模式。
  • 分析集体运动模式及其对 $a$ 的依赖性,包括策略演化中的相变。

提出的方法

  • 在实验室环境中开展控制实验,共84组,每组6名参与者,每人进行300轮广义RPS游戏,采用7种不同的收益矩阵,收益参数 $a$ 不同。
  • 采用随机配对淘汰赛设计,仅提供局部信息,确保参与者仅知晓自身与对手的出招。
  • 通过条件响应行为测量个体策略转换:最优响应(如 $L_-$, $T_+$, $W_0$)与WSLS(如 $L_-$, $L_+$, $W_0$)。
  • 使用非参数相关性分析(Spearman等级相关系数)评估504名参与者-轮次中 $a$ 与行为比例之间的单调关系。
  • 通过转移矢量场映射集体动态,并将实证模式与复制子动力学的理论预测进行比较。
  • 通过比较不同 $a$ 值下的响应模式,分析行为中的相变,尤其关注 $a=2$(中性稳定性)附近的情况。

实验结果

研究问题

  • RQ1获胜收益 $a$ 如何影响石头剪刀布游戏中最优响应行为的频率与结构?
  • RQ2$a$ 如何影响赢则保持输则改变(WSLS)策略的普遍性,特别是左右转向行为之间的平衡?
  • RQ3是否在 $a=2$ 附近出现策略动态的相变,对应于理论上的稳定性阈值?
  • RQ4随着 $a$ 增大,胜负平结果如何差异化地塑造最优响应与WSLS行为?
  • RQ5观察到的行为模式在多大程度上与进化博弈论中标准学习模型的预测一致或偏离?

主要发现

  • 最优响应行为随 $a$ 增加而上升,Spearman等级相关系数为 0.1010(p < 0.0233),表明激励增强时策略响应性显著提高。
  • 作为最优响应组成部分的赢则保持行为随 $a$ 下降(Spearman等级相关系数 = -0.2917,p < 0.0000),而输则左转行为增加(Spearman等级相关系数 = 0.3547,p < 0.0000)。
  • 作为最优响应另一组成部分的平局后右转行为随 $a$ 增加而上升(Spearman等级相关系数 = 0.4317,p < 0.0000),表明对平局的响应方向性更强。
  • 整体WSLS行为随 $a$ 下降(Spearman等级相关系数 = -0.2183,p < 0.0000),尽管输则左转行为增加(rho = 0.3547),而输则右转行为减少(rho = -0.2249)。
  • 在 $a=2$ 附近出现明显的策略动态相变,集体运动从向外发散(离心)矢量场转变为向内汇聚(向心)矢量场,与理论稳定性阈值一致。
  • 随着 $a$ 增大,系统集体策略流动变得更加向心,表明策略分布的稳定性增强,尤其在稳定区间($a > 2$)表现显著。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。