[Paper Review] Brain encoding models based on multimodal transformers can transfer across language and vision
The paper shows that encoding models trained on fMRI responses to language stories can predict brain responses to movies, and vice versa, using BridgeTower multimodal transformers; cross-modality transfer reveals shared semantic representations and that multimodal features outperform unimodal alignments.
Encoding models have been used to assess how the human brain represents concepts in language and vision. While language and vision rely on similar concept representations, current encoding models are typically trained and tested on brain responses to each modality in isolation. Recent advances in multimodal pretraining have produced transformers that can extract aligned representations of concepts in language and vision. In this work, we used representations from multimodal transformers to train encoding models that can transfer across fMRI responses to stories and movies. We found that encoding models trained on brain responses to one modality can successfully predict brain responses to the other modality, particularly in cortical regions that represent conceptual meaning. Further analysis of these encoding models revealed shared semantic dimensions that underlie concept representations in language and vision. Comparing encoding models trained using representations from multimodal and unimodal transformers, we found that multimodal transformers learn more aligned representations of concepts in language and vision. Our results demonstrate how multimodal transformers can provide insights into the brain's capacity for multimodal processing.
Motivation & Objective
- Investigate whether encoding models trained on one modality (language or vision) can predict brain responses to the other modality.
- Determine if multimodal transformer representations align language and vision concepts in the brain.
- Identify semantic dimensions shared across language and vision representations.
- Assess whether multimodal training yields better cross-modal transfer than unimodal feature alignment.
Proposed method
- Use BridgeTower, a multimodal transformer trained on image-text data, to extract stimulus features for stories and movies.
- Train language encoding models on story-fMRI and vision encoding models on movie-fMRI using BridgeTower features.
- Evaluate cross-modality transfer by predicting movie-fMRI from story features and story-fMRI from movie features.
- Align BridgeTower feature spaces with linear mappings estimated from Flickr30K to enable cross-modality projections.
- Conduct voxelwise, L2-regularized regression with hemodynamic delay correction to map stimuli to brain responses.
Experimental results
Research questions
- RQ1Can encoding models trained on language responses predict fMRI responses to visual movie stimuli and vice versa?
- RQ2Do cross-modality transfers reveal aligned semantic representations across language and vision in the cortex?
- RQ3Do multimodal transformer features yield better cross-modality transfer than unimodal features?
- RQ4What semantic dimensions underlie shared language-vision representations in the brain?
Key findings
- Cross-modality encoding performance is positive in many parietal, temporal, and frontal regions outside primary sensory areas.
- Inverted tuning in visual cortex requires correction, which improves cross-modality transfer estimates.
- Cross-modality performance approaches within-modality performance in several regions, indicating similar concept representations across modalities.
- Multimodal BridgeTower features outperform unimodal RoBERTa and ViT features in cross-modality transfer outside visual/auditory cortices.
- PCA on encoding weights reveals shared semantic dimensions across language and vision in multimodal voxels, especially for PCs 1, 3, and 5.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.