Skip to main content
QUICK REVIEW

[Paper Review] OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Hainiu Xu, Runcong Zhao|arXiv (Cornell University)|Feb 8, 2024
Topic Modeling4 citations
TL;DR

OpenToM is a new benchmark for evaluating neural Theory-of-Mind (N-ToM) reasoning in large language models (LLMs), featuring natural, personified narratives with intentional actions, explicit personality traits, and diverse questions targeting both physical-world and psychological mental states. The study reveals that while LLMs perform reasonably on physical-world reasoning, they significantly underperform on psychological-state tracking, exposing critical gaps in their social cognition capabilities.

ABSTRACT

Neural Theory-of-Mind (N-ToM), machine's ability to understand and keep track of the mental states of others, is pivotal in developing socially intelligent agents. However, prevalent N-ToM benchmarks have several shortcomings, including the presence of ambiguous and artificial narratives, absence of personality traits and preferences, a lack of questions addressing characters' psychological mental states, and limited diversity in the questions posed. In response to these issues, we construct OpenToM, a new benchmark for assessing N-ToM with (1) longer and clearer narrative stories, (2) characters with explicit personality traits, (3) actions that are triggered by character intentions, and (4) questions designed to challenge LLMs' capabilities of modeling characters' mental states of both the physical and psychological world. Using OpenToM, we reveal that state-of-the-art LLMs thrive at modeling certain aspects of mental states in the physical world but fall short when tracking characters' mental states in the psychological world.

Motivation & Objective

  • To address shortcomings in existing N-ToM benchmarks, such as artificial narratives, lack of personality, and ambiguous intentions.
  • To develop a comprehensive benchmark that evaluates LLMs' ability to reason about characters' mental states in both the physical and psychological worlds.
  • To generate high-quality, human-verified narratives using a four-stage, LLM-augmented, human-in-the-loop pipeline to reduce spurious cues and improve realism.
  • To assess the performance of state-of-the-art LLMs on psychological reasoning tasks, identifying persistent limitations in N-ToM capabilities.

Proposed method

  • Employ a four-stage human-in-the-loop pipeline: (1) assign personality traits and preferences to characters, (2) define intentions and corresponding actions (enactions), (3) generate naturalistic narratives using LLMs, and (4) refine stories via human annotators for clarity and coherence.
  • Design narratives with two protagonists, an entity-of-interest, and multiple locations, ensuring actions are driven by character intentions and personality.
  • Construct three question types per story: (1) location-based (Loc), (2) multi-hop reasoning (MHop), and (3) attitude-based (Att) targeting psychological states.
  • Use zero-shot evaluation across diverse LLMs (e.g., GPT-4, Llama2, Mixtral) and test prompting techniques like Chain-of-Thought and Simulated-ToM.
  • Fine-tune a Llama2-Chat-13B model to establish a fine-tuning baseline for comparison.
  • Apply rigorous mitigation strategies to reduce spurious correlations, such as lexical overlap, to ensure valid evaluation of true N-ToM reasoning.

Experimental results

Research questions

  • RQ1To what extent can LLMs accurately predict the physical-world location of an entity based on narrative context and character intentions?
  • RQ2How well do LLMs reason about characters’ psychological states, such as attitudes, beliefs, and emotional responses, in complex social narratives?
  • RQ3Does the inclusion of personality traits and intentional actions improve LLMs’ performance on ToM reasoning tasks compared to template-based benchmarks?
  • RQ4How do advanced prompting techniques (e.g., Chain-of-Thought, SimToM) and fine-tuning affect LLMs’ ability to model psychological mental states?
  • RQ5What are the key failure modes of LLMs in N-ToM reasoning, particularly regarding unfaithfulness, sensitivity to narrative length, and role ambiguity?

Key findings

  • State-of-the-art LLMs, including GPT-4 and fine-tuned Llama2-13B, achieve high accuracy on physical-world reasoning (e.g., object location), but performance drops significantly on psychological-state questions.
  • LLMs show a pronounced mismatch in reasoning capabilities: strong on physical-world perception but weak on psychological-state modeling, indicating a fundamental gap in social cognition.
  • Even with advanced prompting (e.g., Chain-of-Thought and SimToM), LLMs fail to consistently deduce characters’ attitudes or beliefs, suggesting inherent limitations in mental state tracking.
  • The models exhibit unfaithfulness in reasoning, often generating plausible but incorrect justifications that do not align with narrative logic or character psychology.
  • Performance is sensitive to narrative length and character roles, with models struggling more when roles are ambiguous or narratives are longer and more complex.
  • The benchmark reveals that current LLMs are prone to relying on spurious cues and fail to understand nuanced psychological perceptions, even in well-constructed, personified stories.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.