Skip to main content
QUICK REVIEW

[Paper Review] How well does surprisal explain N400 amplitude under different experimental conditions?

James A. Michaelov, Benjamin K. Bergen|arXiv (Cornell University)|Oct 9, 2020
Neurobiology of Language and Bilingualism62 references41 citations
TL;DR

This study investigates whether surprisal—measured via recurrent neural network language models—can predict N400 amplitude across diverse neurolinguistic experiments. It finds that surprisal explains N400 amplitude well in most conditions, but fails in cases involving quantifiers, event structure violations, and morphosyntactic anomalies, suggesting that semantic and lexical processing indexed by the N400 involves more than just statistical language prediction.

ABSTRACT

We investigate the extent to which word surprisal can be used to predict a neural measure of human language processing difficulty - the N400. To do this, we use recurrent neural networks to calculate the surprisal of stimuli from previously published neurolinguistic studies of the N400. We find that surprisal can predict N400 amplitude in a wide range of cases, and the cases where it cannot do so provide valuable insight into the neurocognitive processes underlying the response.

Motivation & Objective

  • To test whether surprisal from recurrent neural network language models (RNN-LMs) can predict N400 amplitude across a range of neurolinguistic experimental conditions.
  • To identify systematic discrepancies between surprisal predictions and actual N400 responses, thereby revealing limitations of purely statistical language models in capturing neurocognitive processing.
  • To explore whether additional cognitive mechanisms—such as spreading activation or semantic weighting—beyond bottom-up statistical learning are required to fully explain N400 responses.
  • To evaluate the cognitive plausibility of RNN-LMs as models of human language processing, particularly in relation to the N400 component.

Proposed method

  • Calculated surprisal using two pre-trained RNN-LM architectures (Jozefowicz et al., 2016; Gulordava et al., 2018) on stimuli from 11 experiments across six published studies.
  • Applied the standard surprisal formula: S(wi) = −log P(wi|w1...wi−1), where surprisal reflects the inverse probability of a word given its preceding context.
  • Compared model-predicted surprisal with empirically measured N400 amplitudes from published ERP studies to assess predictive accuracy.
  • Systematically evaluated model performance across varied experimental conditions, including cloze probability, semantic relatedness, typicality, anomaly, quantifiers, and morphosyntactic/event structure violations.
  • Used statistical comparisons to assess whether significant differences in surprisal predicted significant differences in N400 amplitude.
  • Explored potential improvements by considering semantic similarity weighting (e.g., via sentence vectors) to better align model predictions with human neural responses.

Experimental results

Research questions

  • RQ1To what extent does surprisal from RNN-LMs predict N400 amplitude across diverse experimental conditions?
  • RQ2In which linguistic contexts does surprisal fail to predict N400 amplitude, and what do these failures reveal about the underlying neurocognitive mechanisms?
  • RQ3Are there specific linguistic phenomena—such as quantifiers, event structure violations, or morphosyntactic anomalies—where surprisal over- or under-predicts N400 responses?
  • RQ4Can the inclusion of semantic similarity or spreading activation mechanisms improve the predictive power of RNN-LM-based surprisal models for N400 responses?

Key findings

  • Surprisal from RNN-LMs significantly predicts N400 amplitude in 9 out of 11 experimental conditions, including cloze probability, semantic relatedness, and semantic typicality.
  • For stimuli involving quantifiers (Urbach & Kutas, 2010), surprisal failed to replicate the N400 finding that TYPICAL nouns show less amplitude reduction when paired with FEW or RARELY, indicating a need for more explicit representation of quantification.
  • In event structure violation tasks (Kim & Osterhout, 2005), surprisal was more sensitive to violations than N400 amplitude, with attraction violation stimuli showing higher surprisal than controls despite lower N400 amplitudes, suggesting that surprisal overestimates processing difficulty in such cases.
  • For morphosyntactic anomalies (Ainsworth-Darnell et al., 1998), surprisal predicted lower values for semantically anomalous but syntactically acceptable words than for semantically and syntactically anomalous ones, contradicting human N400 data, indicating that models do not fully capture human sensitivity to syntactic structure.
  • The study concludes that while surprisal captures many aspects of N400 responses, it fails in cases requiring deeper semantic integration or syntactic sensitivity, implying that additional mechanisms like spreading activation are needed.
  • The results support the idea that N400 amplitude cannot be fully explained by linguistic input alone and that neurocognitive processes involve more than statistical language prediction.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.