Skip to main content
QUICK REVIEW

[Paper Review] Improving semantic understanding in speech language models via brain-tuning

Omer Moussa, Dietrich Klakow|arXiv (Cornell University)|Oct 11, 2024
Speech and dialogue systems4 citations
TL;DR

This paper introduces 'brain-tuning,' a method that fine-tunes pretrained speech language models using fMRI recordings from individuals listening to natural stories. By incorporating brain activity data during training, the approach significantly improves semantic alignment with human brain regions, reduces reliance on low-level speech features, and boosts performance across diverse downstream semantic tasks—providing the first converging evidence that brain-informed training enhances both brain-like representation and real-world model performance.

ABSTRACT

Speech language models align with human brain responses to natural language to an impressive degree. However, current models rely heavily on low-level speech features, indicating they lack brain-relevant semantics which limits their utility as model organisms of semantic processing in the brain. In this work, we address this limitation by inducing brain-relevant bias directly into the models via fine-tuning with fMRI recordings of people listening to natural stories, a process we name brain-tuning. After testing it on 3 different pretrained model families, we show that brain-tuning not only improves overall alignment with new brain recordings in semantic language regions, but also reduces the reliance on low-level speech features for this alignment. Excitingly, we further show that brain-tuning leads to 1) consistent improvements in performance on a range of downstream tasks and 2) a representational space with increased semantic preference. Our results provide converging evidence, for the first time, that incorporating brain signals into the training of language models improves the models' semantic understanding.

Motivation & Objective

  • To address the limitation that current speech language models rely heavily on low-level speech features rather than brain-relevant semantics, despite strong alignment with brain activity.
  • To develop a method that directly injects brain-relevant semantic bias into pretrained speech models using fMRI recordings.
  • To evaluate whether brain-tuning improves alignment with semantic brain regions, reduces dependence on low-level features, and enhances performance on semantic downstream tasks.
  • To demonstrate that improved brain alignment translates into tangible gains in model utility beyond neuroscience evaluation.

Proposed method

  • Fine-tune three pretrained speech language model families using fMRI recordings collected while participants listened to natural stories.
  • Train the models end-to-end with a loss that minimizes the difference between model activations and fMRI activity in semantic language regions.
  • Use a brain-tuning objective that encourages model representations to align with brain responses to natural language input.
  • Compare brain-tuned models to their pretrained counterparts and two baselines: brain-tuning with block-permuted fMRI data and fine-tuning using larger model representations.
  • Evaluate alignment with new fMRI data in semantic brain regions, assess the influence of low-level features (e.g., triphones, articulation), and measure performance on 5 semantic downstream tasks.
  • Analyze representational similarity and semantic preference in late layers of the models to assess changes in representational space.
(a) Proposed brain-tuning approach
(a) Proposed brain-tuning approach

Experimental results

Research questions

  • RQ1Does brain-tuning improve alignment between speech language models and fMRI recordings in semantic brain regions?
  • RQ2Does brain-tuning reduce the model's reliance on low-level speech features for achieving brain alignment?
  • RQ3Does brain-tuning lead to measurable improvements in downstream tasks requiring semantic understanding?
  • RQ4Does the representational space of brain-tuned models exhibit increased semantic preference?
  • RQ5Is the performance gain from brain-tuning due to the inclusion of brain signals, or could it be replicated by other fine-tuning strategies?

Key findings

  • Brain-tuning significantly improves alignment with new fMRI recordings in semantic language regions across all three tested model families.
  • The method substantially reduces the impact of low-level speech features—such as triphones and articulation—on alignment with semantic brain regions, indicating a shift toward brain-relevant semantics.
  • Brain-tuned models show consistent and significant performance gains on five downstream tasks requiring semantic understanding, including tasks with varying semantic difficulty.
  • Whisper, one of the model families tested, improved from near-baselines to strong performance across all tasks after brain-tuning, indicating broad applicability.
  • The representational space of brain-tuned models exhibits increased semantic preference, particularly in late layers, confirming a shift toward more semantically meaningful representations.
  • The gains were achieved with less than 0.7% of the original training data size, demonstrating high sample efficiency of the brain-tuning approach.
(b) Approach to estimate brain alignment and low-level feature impact
(b) Approach to estimate brain alignment and low-level feature impact

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.