Skip to main content
QUICK REVIEW

[Paper Review] Human-Like Intuitive Behavior and Reasoning Biases Emerged in Language Models -- and Disappeared in GPT-4

Thilo Hagendorff, Sarah Fabi|arXiv (Cornell University)|Jun 13, 2023
Topic Modeling5 citations
TL;DR

This study investigates intuitive reasoning biases in large language models (LLMs) using psychological tests like the Cognitive Reflection Test (CRT) and semantic illusions. It finds that GPT-3 exhibits human-like intuitive errors, while GPT-4 demonstrates hyperrational behavior by avoiding these biases, suggesting a shift from intuitive to analytical reasoning with increased model capability.

ABSTRACT

Large language models (LLMs) are currently at the forefront of intertwining AI systems with human communication and everyday life. Therefore, it is of great importance to evaluate their emerging abilities. In this study, we show that LLMs, most notably GPT-3, exhibit behavior that strikingly resembles human-like intuition -- and the cognitive errors that come with it. However, LLMs with higher cognitive capabilities, in particular ChatGPT and GPT-4, learned to avoid succumbing to these errors and perform in a hyperrational manner. For our experiments, we probe LLMs with the Cognitive Reflection Test (CRT) as well as semantic illusions that were originally designed to investigate intuitive decision-making in humans. Moreover, we probe how sturdy the inclination for intuitive-like decision-making is. Our study demonstrates that investigating LLMs with methods from psychology has the potential to reveal otherwise unknown emergent traits.

Motivation & Objective

  • To investigate whether large language models (LLMs) exhibit human-like intuitive reasoning and associated cognitive biases.
  • To assess how model capabilities influence susceptibility to intuitive errors using established psychological tests.
  • To explore whether advanced LLMs like GPT-4 overcome intuitive biases seen in earlier models such as GPT-3.
  • To evaluate the robustness of intuitive decision-making tendencies in LLMs under varying test conditions.
  • To demonstrate the value of psychological testing methods in uncovering emergent cognitive traits in LLMs.

Proposed method

  • The study employs the Cognitive Reflection Test (CRT), a well-established psychological instrument designed to measure intuitive versus analytical thinking in humans.
  • Semantic illusions—cognitively misleading linguistic stimuli—were used to probe intuitive reasoning in LLMs, similar to human cognitive bias experiments.
  • LLMs including GPT-3, ChatGPT, and GPT-4 were prompted with CRT problems and semantic illusions to assess response patterns.
  • Responses were analyzed for alignment with intuitive (incorrect) answers versus analytical (correct) answers, particularly focusing on error rates.
  • The robustness of intuitive tendencies was tested by varying phrasing and contextual cues to assess consistency of bias across conditions.
  • The study draws comparisons between model outputs and known human cognitive behavior to evaluate the emergence of intuitive-like behavior.

Experimental results

Research questions

  • RQ1Do large language models exhibit human-like intuitive reasoning biases similar to those observed in humans?
  • RQ2How does model capability, particularly in GPT-4, affect susceptibility to cognitive biases like those in the CRT?
  • RQ3To what extent do semantic illusions elicit intuitive errors in LLMs, and how do these compare to human responses?
  • RQ4Are intuitive biases in LLMs robust across different phrasings and contextual manipulations?
  • RQ5Can psychological testing methods reveal emergent cognitive traits in LLMs that are otherwise undetected?

Key findings

  • GPT-3 exhibited a high rate of intuitive errors on the Cognitive Reflection Test, closely mirroring human error patterns.
  • GPT-4 demonstrated a significant reduction in intuitive errors, consistently selecting correct analytical answers over intuitive but incorrect responses.
  • The study found that higher model capabilities correlate with a shift from intuitive to hyperrational decision-making in LLMs.
  • Semantic illusions induced predictable intuitive responses in GPT-3, but GPT-4 largely resisted these biases, indicating stronger analytical reasoning.
  • The robustness of intuitive tendencies was lower in GPT-4, suggesting that its reasoning is less susceptible to linguistic framing effects.
  • The results indicate that psychological testing methods can effectively uncover emergent cognitive traits in LLMs, particularly shifts in reasoning style with model advancement.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.