Skip to main content
QUICK REVIEW

[Paper Review] Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks

Pierre Orhan, Pablo Diego-Simón|arXiv (Cornell University)|Jan 26, 2026
Language Development and Disorders0 citations
TL;DR

The paper shows that self-supervised speech and text models develop phonemic, lexical semantic, and syntactic subspaces in their activations during training, revealed by a shared linear probe, with a sequential emergence and data-demand gap compared to human learning.

ABSTRACT

During language acquisition, children successively learn to categorize phonemes, identify words, and combine them with syntax to form new meaning. While the development of this behavior is well characterized, we still lack a unifying computational framework to explain its underlying neural representations. Here, we investigate whether and when phonemic, lexical, and syntactic representations emerge in the activations of artificial neural networks during their training. Our results show that both speech- and text-based models follow a sequence of learning stages: during training, their neural activations successively build subspaces, where the geometry of the neural activations represents phonemic, lexical, and syntactic structure. While this developmental trajectory qualitatively relates to children's, it is quantitatively different: These algorithms indeed require two to four orders of magnitude more data for these neural representations to emerge. Together, these results show conditions under which major stages of language acquisition spontaneously emerge, and hence delineate a promising path to understand the computations underpinning language acquisition.

Motivation & Objective

  • Motivate a unified computational framework to explain neural representations underlying language acquisition.
  • Investigate whether phonemic, lexical semantic, and syntactic representations emerge in neural activations during training.
  • Characterize the geometry and order of emergence of these linguistic structures across modalities and models.
  • Assess data efficiency and how emergence in models compares to human language acquisition.

Proposed method

  • Generalize Hewitt and Manning (2019) Structural Probe to extract phonemic, lexical semantic, and syntactic subspaces from model activations.
  • Fit linear transformations B (2D for visualization, 200D for evaluation) to align activation distances with linguistic target distances.
  • Evaluate probe performance via Spearman correlation between target distances and projected distances across phoneme, lexical, and syntax levels.
  • Construct probing datasets: UD-EWT for syntax, WordNet nouns for lexical semantics, and phoneme-based representations derived from TTS-synthesized speech with alignments.
  • Compare text models (Pythia, Llama2) and speech models (Wav2Vec 2.0) across model sizes and pretraining conditions.
  • Assess emergence by tracking probe scores across training checkpoints and pretraining steps.

Experimental results

Research questions

  • RQ1Do phonemic, lexical semantic, and syntactic structures emerge as separable subspaces in neural activations of speech and text models?
  • RQ2What is the order of emergence of these linguistic representations during training, and how does data quantity affect it?
  • RQ3How does model type (text vs. speech) and model size influence the emergence and geometry of these structures?
  • RQ4To what extent do acoustic cues confound semantic representations in audio models, and how do control conditions address this?
  • RQ5Are the findings consistent with a developmental trajectory analogous to human language acquisition?

Key findings

  • Phonemic structure is recoverable as a distinct subspace in speech models, with articulation-like geometry (e.g., vowel relationships) emerging in mid-to-late layers during pretraining.
  • Lexical semantic structure shows detectable, but more modest, organization in both text and audio models, highly dependent on model size and data exposure.
  • Syntactic representations are recoverable in both speech and text models, with strong scores that plateau with model size but faster emergence in audio models due to cues in speech data.
  • Across checkpoints, phonemic emergence precedes partial lexical-semantic emergence, which in turn precedes syntactic emergence, indicating a sequential developmental trajectory.
  • Audio models require substantially more input data to reach comparable representations than human children, revealing a data-efficiency gap.
  • Controls show semantic and syntactic structures in audio models are not solely attributable to acoustic cues; text models demonstrate stronger and clearer semantic/syntactic structuring.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.