Skip to main content
QUICK REVIEW

[Paper Review] A fine-grained comparison of pragmatic language understanding in humans and language models

Jennifer J. Hu, Sammy Floyd|arXiv (Cornell University)|Dec 13, 2022
Topic Modeling4 citations
TL;DR

This study conducts a fine-grained comparison of human and language model (LM) performance on seven pragmatic language phenomena using zero-shot prompting on expert-curated English materials. Large models like Flan-T5 (XL) and text-davinci-002 achieve high accuracy and mirror human error patterns—favoring literal interpretations over heuristic-based distractors—suggesting pragmatic behaviors can emerge without explicit mental state representations, though models struggle with irony, humor, and social expectation violations.

ABSTRACT

Pragmatics and non-literal language understanding are essential to human communication, and present a long-standing challenge for artificial language models. We perform a fine-grained comparison of language models and humans on seven pragmatic phenomena, using zero-shot prompting on an expert-curated set of English materials. We ask whether models (1) select pragmatic interpretations of speaker utterances, (2) make similar error patterns as humans, and (3) use similar linguistic cues as humans to solve the tasks. We find that the largest models achieve high accuracy and match human error patterns: within incorrect responses, models favor literal interpretations over heuristic-based distractors. We also find preliminary evidence that models and humans are sensitive to similar linguistic cues. Our results suggest that pragmatic behaviors can emerge in models without explicitly constructed representations of mental states. However, models tend to struggle with phenomena relying on social expectation violations.

Motivation & Objective

  • To assess whether large language models can recover pragmatic interpretations of utterances without fine-tuning, using zero-shot prompting.
  • To investigate whether language models make similar errors as humans when failing to select the correct pragmatic interpretation.
  • To examine whether models and humans rely on similar linguistic cues to resolve pragmatic ambiguity.
  • To explore whether pragmatic competence in models emerges without explicitly constructed representations of mental states.
  • To identify model weaknesses in understanding phenomena involving social norms, expectations, and irony.

Proposed method

  • Use of zero-shot prompting to evaluate language models on a curated set of 72 multiple-choice questions covering diverse pragmatic phenomena.
  • Employment of four distinct language models: GPT-2, T5-Instruct, Flan-T5 (XL), and InstructGPT (text-davinci-002), representing varying sizes and training objectives.
  • Collection of human response data from expert-validated stimuli to establish baseline error patterns and cue sensitivity.
  • Fine-grained analysis of model responses, comparing selection distributions, error types (literal vs. heuristic-based), and cue sensitivity to human patterns.
  • Use of a manually curated dataset designed by expert researchers to ensure linguistic and pragmatic validity across phenomena including irony, indirect requests, and scalar implicatures.
  • Comparison of model and human response distributions using statistical measures to assess similarity in correct responses and error patterns.

Experimental results

Research questions

  • RQ1Do language models recover the intended pragmatic interpretation of speaker utterances in zero-shot settings?
  • RQ2When language models fail, do they make the same types of errors as humans—particularly favoring literal interpretations over heuristic-based distractors?
  • RQ3Are language models and humans sensitive to the same linguistic cues when resolving pragmatic ambiguity?
  • RQ4Can pragmatic behaviors emerge in language models without explicit representations of agents’ mental states?
  • RQ5How do models perform on pragmatic phenomena that rely on social expectation violations, such as irony and humor?

Key findings

  • The largest models, particularly Flan-T5 (XL) and OpenAI’s text-davinci-002, achieve high accuracy on pragmatic understanding tasks, with performance approaching that of humans.
  • When incorrect, models predominantly select the literal interpretation rather than heuristic-based distractors, mirroring the error pattern observed in human participants.
  • There is preliminary evidence that models and humans are sensitive to similar linguistic cues in resolving pragmatic ambiguity.
  • Models show particular difficulty with phenomena involving social expectation violations, such as irony, humor, and conversational maxims.
  • The results suggest that pragmatic competence can emerge in language models without explicitly constructed representations of mental states or belief updates.
  • Despite high accuracy, models remain unpredictable in real-world deployment, especially for API-based models like text-davinci-002, due to limited transparency in training protocols.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.