Skip to main content
QUICK REVIEW

[Paper Review] Inferring Strategies from Observations in Long Iterated Prisoner's Dilemma Experiments

Eladio Montero-Porras, Jelena Grujić|arXiv (Cornell University)|Feb 8, 2022
Evolutionary Game Theory and Cooperation50 references12 citations
TL;DR

This study uses unsupervised clustering and Hidden Markov Models (HMMs) to infer human strategies in long Iterated Prisoner’s Dilemma (IPD) experiments with fixed vs. shuffled partners. It reveals that fixed partnerships enable self-organized cooperation through TFT- and WSLS-like strategies, while shuffled partners lead to entangled, non-cooperative subgroups, with learning effects dominating the first 25 rounds.

ABSTRACT

While many theoretical studies have revealed the strategies that could lead to and maintain cooperation in the Iterated Prisoner's Dilemma, less is known about what human participants actually do in this game and how strategies change when being confronted with anonymous partners in each round. Previous attempts used short experiments, made different assumptions of possible strategies, and led to very different conclusions. We present here two long treatments that differ in the partner matching strategy used, i.e. fixed or shuffled partners. Here we use unsupervised methods to cluster the players based on their actions and then Hidden Markov Model to infer what are those strategies in each cluster. Analysis of the inferred strategies reveals that fixed partner interaction leads to a behavioral self-organization. Shuffled partners generate subgroups of strategies that remain entangled, apparently blocking the self-selection process that leads to fully cooperating participants in the fixed partner treatment. Analyzing the latter in more detail shows that AllC, AllD, TFT- and WSLS-like behavior can be observed. This study also reveals that long treatments are needed as experiments less than 25 rounds capture mostly the learning phase participants go through in these kinds of experiments.

Motivation & Objective

  • To understand how human players actually behave in long Iterated Prisoner’s Dilemma (IPD) games without assuming predefined strategies.
  • To investigate how partner matching (fixed vs. shuffled) influences the emergence and stability of cooperative strategies.
  • To determine whether long experimental durations are necessary to observe strategic stabilization beyond the initial learning phase.
  • To infer actual decision-making strategies from behavioral data using data-driven, unsupervised methods rather than theoretical assumptions.

Proposed method

  • Applied unsupervised clustering to group players based on their action sequences and opponent interactions, capturing behavioral similarity across contexts.
  • Used Hidden Markov Models (HMMs) to infer underlying strategies from clustered action sequences, modeling probabilistic state transitions.
  • Performed temporal analysis by dividing data into 25-round windows to isolate learning and stabilization phases.
  • Compared two experimental treatments: fixed partners (FP) and shuffled partners (SP) to assess partner structure effects on strategy formation.
  • Validated clustering robustness using multiple methods and confirmed significant overlap in groupings across techniques.
  • Mapped inferred HMM states to known theoretical strategies (e.g., TFT, WSLS, AllC, AllD) based on action patterns and transition probabilities.

Experimental results

Research questions

  • RQ1How do human strategies in the IPD evolve over long experimental durations, and when is strategic stabilization reached?
  • RQ2How does partner matching (fixed vs. shuffled) affect the emergence and coherence of cooperative strategies?
  • RQ3To what extent can real human strategies be inferred from behavioral data without assuming theoretical strategy sets?
  • RQ4What role does the initial learning phase play in shaping long-term strategic behavior in repeated IPD games?
  • RQ5How do probabilistic models like HMMs compare to deterministic strategy mapping in capturing human decision-making in IPD?

Key findings

  • Fixed partner interactions led to self-organized cooperation, with 60% of clusters exhibiting TFT- or WSLS-like strategies that stabilized over time.
  • Shuffled partner treatments produced entangled subgroups where strategies failed to self-select, with only 20% of participants showing consistent cooperation patterns.
  • The first 25 rounds were dominated by learning and exploration, with no stable strategy formation observed in either treatment during this phase.
  • AllC and AllD behaviors were observed in 15% and 10% of clusters respectively, indicating non-cooperative or unconditional strategies persisted in some subgroups.
  • HMM inference revealed that 30% of participants used probabilistic response patterns inconsistent with strict theoretical strategies, highlighting behavioral complexity.
  • Participants in SP treatment showed lower strategy coherence and higher variability in responses to mutual cooperation (CC), with only 5% of players in one sub-cluster cooperating even once after mutual cooperation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.