[Paper Review] Evaluating Large Language Models in Theory of Mind Tasks
The study evaluates 11 LLMs on a battery of 640 false-belief prompts across 40 tasks; GPT-4 solves 75% of tasks, approaching six-year-old child performance, while smaller models solve none and earlier models reach ~20%.
Eleven Large Language Models (LLMs) were assessed using a custom-made battery of false-belief tasks, considered a gold standard in testing Theory of Mind (ToM) in humans. The battery included 640 prompts spread across 40 diverse tasks, each one including a false-belief scenario, three closely matched true-belief control scenarios, and the reversed versions of all four. To solve a single task, a model needed to correctly answer 16 prompts across all eight scenarios. Smaller and older models solved no tasks; GPT-3-davinci-003 (from November 2022) and ChatGPT-3.5-turbo (from March 2023) solved 20% of the tasks; ChatGPT-4 (from June 2023) solved 75% of the tasks, matching the performance of six-year-old children observed in past studies. We explore the potential interpretation of these findings, including the intriguing possibility that ToM, previously considered exclusive to humans, may have spontaneously emerged as a byproduct of LLMs' improving language skills.
Motivation & Objective
- Assess whether large language models exhibit theory-of-mind (ToM) abilities in a battery of false-belief tasks.
- Compare performance across eleven LLMs from small/older to state-of-the-art models.
- Quantify the emergence of ToM-like skills as models improve language capabilities.
Proposed method
- Use a custom battery of false-belief tasks (640 prompts across 40 tasks).
- Each task includes a false-belief scenario, three matched true-belief controls, and reversed versions of all four scenarios.
- A model must answer 16 prompts correctly across eight scenarios to solve a single task.
- Evaluation spans eleven LLMs with varying sizes and architectures.
- Results discuss performance relative to human benchmarks from prior ToM studies.
- Code and tasks are made available for replication (Colab).
Experimental results
Research questions
- RQ1How do eleven LLMs perform on a battery of theory-of-mind false-belief tasks?
- RQ2Does ToM performance improve with newer, more capable language models?
- RQ3Are ToM-like capabilities emergent as a byproduct of language modeling progress?
- RQ4How does model performance compare to human ToM benchmarks from prior research?
Key findings
- Smaller and older LLMs solved no tasks in the battery.
- GPT-3-davinci-003 (Nov 2022) and ChatGPT-3.5-turbo (Mar 2023) solved about 20% of the tasks.
- ChatGPT-4 (Jun 2023) solved about 75% of the tasks.
- GPT-4’s performance matched the level observed in past studies of six-year-old children.
- The results raise the possibility that ToM may emerge spontaneously as language skills improve in LLMs.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.