Skip to main content
QUICK REVIEW

[論文レビュー] Incentive and stability in the Rock-Paper-Scissors game: an experimental investigation

Zhijian Wang, Bin Xu|arXiv (Cornell University)|Jul 4, 2014
Evolutionary Game Theory and Cooperation参考文献 56被引用数 5
ひとこと要約

本実験的研究では、一般化されたグーチョキパー・ゲームにおける報酬水準(勝利報酬 $a$)が個人的および集団的戦略ダイナミクスに与える影響を調査している。84組のグループで720ラウンドの対戦を実施した結果、$a$ の増加に伴い最適応答行動が増加し、勝ち続け・負けたら変更する(WSLS)行動が減少することが判明。$a=2$ の周辺で顕著な段階的転移が観察され、報酬構造に応じた人間の学習戦略の体系的変化が明らかになった。

ABSTRACT

In a two-person Rock-Paper-Scissors (RPS) game, if we set a loss worth nothing and a tie worth 1, and the payoff of winning (the incentive a) as a variable, this game is called as generalized RPS game. The generalized RPS game is a representative mathematical model to illustrate the game dynamics, appearing widely in textbook. However, how actual motions in these games depend on the incentive has never been reported quantitatively. Using the data from 7 games with different incentives, including 84 groups of 6 subjects playing the game in 300-round, with random-pair tournaments and local information recorded, we find that, both on social and individual level, the actual motions are changing continuously with the incentive. More expressively, some representative findings are, (1) in social collective strategy transit views, the forward transition vector field is more and more centripetal as the stability of the system increasing; (2) In the individual behavior of strategy transit view, there exists a phase transformation as the stability of the systems increasing, and the phase transformation point being near the standard RPS; (3) Conditional response behaviors are structurally changing accompanied by the controlled incentive. As a whole, the best response behavior increases and the win-stay lose-shift (WSLS) behavior declines with the incentive. Further, the outcome of win, tie, and lose influence the best response behavior and WSLS behavior. Both as the best response behavior, the win-stay behavior declines with the incentive while the lose-left-shift behavior increase with the incentive. And both as the WSLS behavior, the lose-left-shift behavior increase with the incentive, but the lose-right-shift behaviors declines with the incentive. We hope to learn which one in tens of learning models can interpret the empirical observation above.

研究の動機と目的

  • 一般化されたグーチョキパー・ゲームにおける戦略ダイナミクスに、勝利報酬 $a$ の変化が与える影響を実証的に検討すること。
  • 個人の学習行動(特に最適応答とWSLS)が報酬水準に応じて体系的に変化するかどうかを調査すること。
  • 理論的安定性閾値 $a=2$ の周辺で、条件付き応答行動に構造的変化が生じるかどうかを同定すること。
  • 既存の学習モデルが、変化する報酬下での観察された行動パターンを説明できるかどうかを検証すること。
  • 集団的運動パターンとその $a$ 依存性を分析し、戦略進化における段階的転移を特定すること。

提案手法

  • 84組のグループ(各グループ6名)を対象に、7種類の異なる報酬行列($a$ の値が異なる)を用いた制御実験を実施。300ラウンドの一般化グーチョキパーを実施。
  • ローカル情報に基づくランダムペアトーナメント設計を採用。参加者は自分自身と相手の手のみを把握可能。
  • 条件付き応答行動(最適応答:例 $L_-$, $T_+$, $W_0$;WSLS:例 $L_-$, $L_+$, $W_0$)を用いて、個人の戦略転換を測定。
  • 504件の被験者・ラウンド観察データを用い、$a$ と行動割合の間の単調関係を評価する非パラメトリック相関(スピアマンのrho)を適用。
  • 遷移ベクトル場を用いて集団的ダイナミクスをマッピングし、理論的予測(リプロダクター・ダイナミクス)と照合。
  • 特に $a=2$(中立的安定性)の周辺で、$a$ の値ごとの応答パターンを比較し、行動の段階的転移を分析。

実験結果

リサーチクエスチョン

  • RQ1勝利報酬 $a$ がグーチョキパーにおける最適応答行動の頻度と構造に与える影響は何か?
  • RQ2報酬 $a$ が勝ち続け・負けたら変更する(WSLS)戦略の普及に与える影響、特に左シフトと右シフト行動のバランスに与える影響は何か?
  • RQ3理論的安定性閾値に対応する $a=2$ の周辺で、戦略ダイナミクスに段階的転移が生じるか?
  • RQ4勝利、引き分け、敗北の結果が、$a$ の増加に伴い最適応答行動とWSLS行動にどのように異なる影響を与えるか?
  • RQ5観察された行動パターンが、進化的ゲーム理論における標準的学習モデルの予測とどの程度一致するか、あるいは逸脱するか?

主な発見

  • 最適応答行動は $a$ が増加するにつれて増加し、スパイアマンのrho = 0.1010(p < 0.0233)となり、報酬の増大に伴い戦略的反応性が統計的に有意に上昇することが示された。
  • 最適応答の一部である勝ち続け行動は $a$ が増加するにつれて減少(rho = -0.2917, p < 0.0000)し、負けたら左にシフトする行動は増加(rho = 0.3547, p < 0.0000)した。
  • 別の最適応答要因である引き分け後に右にシフトする行動も $a$ が増加するにつれて増加(rho = 0.4317, p < 0.0000)し、引き分けに対するより強い方向性のシフトが確認された。
  • WSLS行動全体としての割合は $a$ が増加するにつれて減少(rho = -0.2183, p < 0.0000)したが、負けたら左にシフトする行動は増加(rho = 0.3547)し、負けたら右にシフトする行動は減少(rho = -0.2249)した。
  • 明確な段階的転移が $a=2$ の周辺に観察され、集団的運動が外向き(発散的)から内向き(中心指向)のベクトル場にシフトした。これは理論的安定性閾値と一致した。
  • $a$ が増加するにつれて、集団戦略の流れがより中心指向的(centripetal)となり、特に $a > 2$ の安定領域において、戦略分布の安定性が高まっていることが示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。