Skip to main content
QUICK REVIEW

[論文レビュー] On the Emergence of Cooperation in the Repeated Prisoner's Dilemma

Maximilian Schaefer|arXiv (Cornell University)|Nov 24, 2022
Evolutionary Game Theory and Cooperation被引用数 5
ひとこと要約

本稿では、1期記憶を持つ$ε$-greedy Q-learnerのシミュレーションを用いて、グリム・トリガー戦略下での確率的リプロダクター動的におけるポテンシャル関数が、繰り返し繰り返しのジレンマにおける協力の出現を予測することを示している。主な結果は、協力的・非協力的パラメータ領域を分ける臨界の運動エネルギー比$\mathcal{C} = \frac{\delta}{1-\delta}\mathcal{K}(\alpha)\epsilon$であり、このフロンティアは人間の協力率を非常に良く予測しており、ピアソン積率相関係数が0.8以上に達する。

ABSTRACT

Using simulations between pairs of $ε$-greedy q-learners with one-period memory, this article demonstrates that the potential function of the stochastic replicator dynamics (Foster and Young, 1990) allows it to predict the emergence of error-proof cooperative strategies from the underlying parameters of the repeated prisoner's dilemma. The observed cooperation rates between q-learners are related to the ratio between the kinetic energy exerted by the polar attractors of the replicator dynamics under the grim trigger strategy. The frontier separating the parameter space conducive to cooperation from the parameter space dominated by defection can be found by setting the kinetic energy ratio equal to a critical value, which is a function of the discount factor, $f(δ) = δ/(1-δ)$, multiplied by a correction term to account for the effect of the algorithms' exploration probability. The gradient at the frontier increases with the distance between the game parameters and the hyperplane that characterizes the incentive compatibility constraint for cooperation under grim trigger. Building on literature from the neurosciences, which suggests that reinforcement learning is useful to understanding human behavior in risky environments, the article further explores the extent to which the frontier derived for q-learners also explains the emergence of cooperation between humans. Using metadata from laboratory experiments that analyze human choices in the infinitely repeated prisoner's dilemma, the cooperation rates between humans are compared to those observed between q-learners under similar conditions. The correlation coefficients between the cooperation rates observed for humans and those observed for q-learners are consistently above $0.8$. The frontier derived from the simulations between q-learners is also found to predict the emergence of cooperation between humans.

研究の動機と目的

  • 1期記憶を持つQ-learnerが繰り返しのジレンマで協力を学ぶ条件を同定すること。
  • 進化的ゲーム理論の概念、特にグリム・トリガー戦略下での確率的リプロダクター動的のポテンシャル関数を用いて、協力の出現を予測すること。
  • Q-learnerのシミュレーションから導かれた協力フロンティアが、実験室での人間の協力行動を予測できるかどうかを検証すること。
  • ゲームパラメータ、学習アルゴリズムのハイパーパrameter、観察された協力率との関係を定量化すること。
  • 異なるQ-learnerの初期化スケームとパラメータ範囲において、協力フロンティアの頑健性を評価すること。

提案手法

  • 繰り返しのジレンマにおいて、一定の学習率($\alpha \in [0.01, 0.1]$)と探索確率($\epsilon \in [0.01, 0.1]$)を持つ$ε$-greedy Q-learnerのペアをシミュレーションする。
  • FosterとYoung(1990)が提唱した、確率的リプロダクター動的のポテンシャル関数を用い、相互協力および相互裏切りに向けた吸引する運動エネルギーを計算する。
  • 割引率$\delta$と探索効果補正係数$\mathcal{K}(\alpha)$を用いた臨界運動エネルギー比$\mathcal{C} = \frac{\delta}{1-\delta}\mathcal{K}(\alpha)\epsilon$を定義する。
  • 運動エネルギー比が$\mathcal{C}$に等しいパラメータの組合せとして構成される協力フロンティアを、協力的・非協力的領域を分ける境界として特定する。
  • Q-learnerの協力率と実験室実験における人間の協力率を、$(d^{ic}, \mathcal{KLR} - \log(\mathcal{K}(\alpha)\epsilon))$空間におけるユークリッド距離で比較する。
  • ゲーム経験の増加に伴い、人間とQ-learnerの協力率のピアソン積率相関係数を算出し、変化を追跡する。

実験結果

リサーチクエスチョン

  • RQ1繰り返しのジレンマにおいて、1期記憶を持つQ-learnerが安定した協力を示すパラメータの組合せは何か?
  • RQ2グリム・トリガー戦略下での確率的リプロダクター動的のポテンシャル関数は、Q-learner相互作用における協力的・非協力的結果の境界を予測できるか?
  • RQ3Q-learnerのシミュレーションから導かれた協力フロンティアは、実験室実験における実際の人間の協力率をどの程度正確に予測できるか?
  • RQ4ゲームパラメータとインcentive compatibility制約超平面との距離に応じて、フロンティア付近での協力率勾配はどのように変化するか?
  • RQ5プレイヤーが繰り返しゲームで経験を積むに従い、人間とQ-learnerの協力率の相関はどのように変化するか?

主な発見

  • Q-learner間の協力フロンティアは、探索効果補正を加えた運動エネルギー比$\mathcal{C} = \frac{\delta}{1-\delta}\mathcal{K}(\alpha)\epsilon$によって正確に予測可能であり、$\mathcal{K}(\alpha)$は探索効果を補正する。
  • フロンティアにおける協力戦略の割合の勾配は、ゲームパラメータとインcentive compatibility超平面との距離が増加するに従い上昇し、その距離が最大値の50%を超えると安定化する。
  • テスト範囲内において、協力フロンティアは楽観的および悲観的なQ値初期化の両方に対して頑健である。
  • 実験室実験における人間の協力率は、Q-learnerの協力率とピアソン積率相関係数0.8以上を示しており、Q-learnerフロンティアの強い予測力が裏付けられる。
  • 7ゲーム目を過ぎても相関係数は0.8以上を維持し、早期にピークに達し、安定化する傾向を示しており、人間とQ-learnerの学習ダイナミクスの整合性が示唆される。
  • 唯一、$\mathcal{KLR}$と$sizeGOOD$指標で予測が食い違う処理では、$\mathcal{KLR}$フロンティアが妥当であることを裏付ける結果が得られた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。