Skip to main content
QUICK REVIEW

[Paper Review] Strategic Behavior of Large Language Models: Game Structure vs. Contextual Framing

Nunzio Lorè, Babak Heydari|arXiv (Cornell University)|Sep 12, 2023
Topic ModelingComputer Science3 citations
TL;DR

This study evaluates the strategic decision-making of GPT-3.5, GPT-4, and LLaMa-2 in four two-player social dilemma games—Prisoner’s Dilemma, Stag Hunt, Snowdrift, and Prisoner’s Delight—under varying contextual framings (e.g., diplomatic, friendly, business). Results show that while all models are influenced by context, LLaMa-2 demonstrates superior strategic reasoning by distinguishing game structures more accurately than GPT-4, which relies on a binary categorization of games, and GPT-3.5, which is highly context-sensitive but lacks abstract strategic reasoning.

ABSTRACT

This paper investigates the strategic decision-making capabilities of three Large Language Models (LLMs): GPT-3.5, GPT-4, and LLaMa-2, within the framework of game theory. Utilizing four canonical two-player games -- Prisoner's Dilemma, Stag Hunt, Snowdrift, and Prisoner's Delight -- we explore how these models navigate social dilemmas, situations where players can either cooperate for a collective benefit or defect for individual gain. Crucially, we extend our analysis to examine the role of contextual framing, such as diplomatic relations or casual friendships, in shaping the models' decisions. Our findings reveal a complex landscape: while GPT-3.5 is highly sensitive to contextual framing, it shows limited ability to engage in abstract strategic reasoning. Both GPT-4 and LLaMa-2 adjust their strategies based on game structure and context, but LLaMa-2 exhibits a more nuanced understanding of the games' underlying mechanics. These results highlight the current limitations and varied proficiencies of LLMs in strategic decision-making, cautioning against their unqualified use in tasks requiring complex strategic reasoning.

Motivation & Objective

  • To investigate how Large Language Models (LLMs) make strategic decisions in canonical two-player games involving cooperation and defection.
  • To assess the relative influence of game structure versus contextual framing—such as diplomatic, business, or friendly interactions—on LLM decision-making.
  • To compare the strategic reasoning capabilities of three prominent LLMs: GPT-3.5, GPT-4, and LLaMa-2.
  • To evaluate whether LLMs can perform abstract strategic reasoning or are primarily reactive to contextual cues.
  • To explore the implications of these findings for the use of LLMs in real-world strategic or social decision-making scenarios.

Proposed method

  • Conducted game-theoretic simulations using four canonical two-player games: Prisoner’s Dilemma, Stag Hunt, Snowdrift, and Prisoner’s Delight.
  • Applied five distinct contextual framings—diplomatic, business, friendly, casual, and adversarial—to each game to assess framing effects.
  • Collected and analyzed model responses to determine strategic choices (cooperate or defect) and the reasoning behind them.
  • Used qualitative analysis of reasoning traces to evaluate depth of strategic understanding, including identification of best response strategies.
  • Compared model behavior across game types and contexts to assess sensitivity to structural game features versus contextual cues.
  • Employed anecdotal and pattern-based analysis of reasoning chains to infer underlying decision mechanisms, such as reliance on payoff maximization or context-driven norms.
Figure 1 : A schematic explanation of our data collecting process. A combination of a contextual prompt and a game prompt is fed into one of the LLM we examine in this paper, namely GPT-3.5, GPT-4, and LLaMa-2. Each combination creates a unique scenario, and for each scenario we collect 300 initiali
Figure 1 : A schematic explanation of our data collecting process. A combination of a contextual prompt and a game prompt is fed into one of the LLM we examine in this paper, namely GPT-3.5, GPT-4, and LLaMa-2. Each combination creates a unique scenario, and for each scenario we collect 300 initiali

Experimental results

Research questions

  • RQ1How do GPT-3.5, GPT-4, and LLaMa-2 differ in their strategic choices across four canonical two-player games?
  • RQ2To what extent does contextual framing—such as friendship or business relationships—affect the strategic decisions of LLMs?
  • RQ3How well do LLMs recognize and reason about the structural differences between game types like Stag Hunt and Snowdrift?
  • RQ4Does GPT-4’s prioritization of game structure lead to nuanced strategic reasoning, or does it rely on a simplified binary classification of games?
  • RQ5Can LLaMa-2’s stronger performance in distinguishing game structures be attributed to a more refined understanding of strategic incentives?

Key findings

  • GPT-3.5 is highly sensitive to contextual framing but shows limited ability to perform abstract strategic reasoning, often failing to identify best response strategies.
  • GPT-4 prioritizes game structure over context but applies a binary classification of games into 'high' or 'low' social dilemma categories, lacking nuanced differentiation between distinct game types.
  • LLaMa-2 demonstrates a more finely-grained understanding of game structures compared to GPT-4, even though it places greater emphasis on contextual factors.
  • In games like Prisoner’s Delight, LLaMa-2 consistently reasons toward the individually and socially optimal choice (cooperation) when reasoning is correct, though it occasionally makes basic mathematical errors.
  • GPT-4 often mischaracterizes non-Prisoner’s Dilemma games as variants of the Prisoner’s Dilemma, indicating a lack of structural discrimination.
  • Contextual framing significantly influences decisions, especially in friendly or collaborative contexts, where even structurally non-cooperative games may elicit cooperative responses, suggesting susceptibility to framing manipulation.
(a) Results grouped game, GPT-3.5
(a) Results grouped game, GPT-3.5

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.