Skip to main content
QUICK REVIEW

[Paper Review] Are ChatGPT and GPT-4 Good Poker Players? -- A Pre-Flop Analysis

Akshat Gupta|arXiv (Cornell University)|Aug 23, 2023
Sports Analytics and Performance4 citations
TL;DR

This paper evaluates ChatGPT and GPT-4 as pre-flop Texas No-Limit Hold’em poker players using game theory optimal (GTO) benchmarks. Despite advanced understanding of poker concepts like hand ranges and position, both models fail to achieve GTO play: ChatGPT plays too conservatively (a 'nit'), while GPT-4 is overly aggressive (a 'maniac'), indicating fundamental misalignment with optimal strategy despite strong domain comprehension.

ABSTRACT

Since the introduction of ChatGPT and GPT-4, these models have been tested across a large number of tasks. Their adeptness across domains is evident, but their aptitude in playing games, and specifically their aptitude in the realm of poker has remained unexplored. Poker is a game that requires decision making under uncertainty and incomplete information. In this paper, we put ChatGPT and GPT-4 through the poker test and evaluate their poker skills. Our findings reveal that while both models display an advanced understanding of poker, encompassing concepts like the valuation of starting hands, playing positions and other intricacies of game theory optimal (GTO) poker, both ChatGPT and GPT-4 are NOT game theory optimal poker players. Profitable strategies in poker are evaluated in expectations over large samples. Through a series of experiments, we first discover the characteristics of optimal prompts and model parameters for playing poker with these models. Our observations then unveil the distinct playing personas of the two models. We first conclude that GPT-4 is a more advanced poker player than ChatGPT. This exploration then sheds light on the divergent poker tactics of the two models: ChatGPT's conservativeness juxtaposed against GPT-4's aggression. In poker vernacular, when tasked to play GTO poker, ChatGPT plays like a nit, which means that it has a propensity to only engage with premium hands and folds a majority of hands. When subjected to the same directive, GPT-4 plays like a maniac, showcasing a loose and aggressive style of play. Both strategies, although relatively advanced, are not game theory optimal.

Motivation & Objective

  • To assess whether large language models like ChatGPT and GPT-4 can play Texas No-Limit Hold’em poker at a game theory optimal (GTO) level.
  • To investigate how these models interpret and apply GTO poker principles when prompted.
  • To compare the strategic behaviors of ChatGPT and GPT-4 in pre-flop decision-making under incomplete information.
  • To identify the root causes of deviation from GTO play in LLMs, particularly in aggression and hand-range selection.
  • To explore the implications of LLMs' strategic misalignment for their use in complex, uncertainty-driven decision-making tasks.

Proposed method

  • The study evaluates pre-flop decisions in a 9-player Texas No-Limit Hold’em poker game, focusing on the raise-first-in scenario.
  • Models are prompted with standardized instructions: basic play, GTO play, and role-specific prompts to elicit strategic behavior.
  • A total of hundreds of thousands of API queries are made to assess model responses across diverse starting hands and positions.
  • Responses are analyzed by mapping them to pre-defined GTO hand ranges (e.g., from GTO solvers like GUNHOE) to measure deviation.
  • Statistical comparison is performed between model actions and optimal GTO strategies, particularly in terms of limping, raising, and folding frequencies.
  • The analysis focuses on position-dependent behavior, especially in early, middle, and late positions (UTG, MP, CO, BTN).

Experimental results

Research questions

  • RQ1To what extent do ChatGPT and GPT-4 align with GTO pre-flop strategies in Texas No-Limit Hold’em?
  • RQ2How do the strategic behaviors of ChatGPT and GPT-4 differ when prompted to play GTO poker versus their default behavior?
  • RQ3Why do both models fail to achieve GTO play despite demonstrating strong understanding of poker concepts like position and hand strength?
  • RQ4What role does model architecture (e.g., GPT-4 vs. ChatGPT) play in shaping aggressive or conservative playing styles?
  • RQ5How do prompt engineering and model identity influence the deviation from optimal strategy in incomplete information games?

Key findings

  • ChatGPT exhibits a conservative playing style, folding the majority of hands and only playing with premium starting hands, resembling a 'nit' in poker terminology.
  • GPT-4 demonstrates a highly aggressive style, raising a significantly larger proportion of hands than optimal, especially from late positions like the Button.
  • When prompted to play GTO, GPT-4 increases its raising frequency even further, particularly from the Button, where it raises 90% of hands—well above the GTO benchmark.
  • Despite its advanced understanding of poker rules and concepts, GPT-4 never limps in any scenario, indicating a fundamental bias toward aggression.
  • ChatGPT becomes more aggressive when prompted to play GTO, but still folds too many hands, indicating a persistent conservative bias.
  • Both models fail to achieve GTO play: ChatGPT due to insufficient aggression, GPT-4 due to excessive aggression, revealing a lack of self-awareness about their own strategic deviations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.