[论文解读] Visio-Linguistic Brain Encoding
本文利用视觉-语言Transformer(如VisualBERT、LXMERT)研究视觉-语言脑编码,旨在从视觉和语言刺激预测fMRI脑活动。结果表明,多模态Transformer在BOLD5000和Pereira数据集上的表现显著优于单模态模型(CNN、图像Transformer),达到新的最先进水平,提示语言处理可能在被动观看图像时隐式影响视觉脑区响应。
Enabling effective brain-computer interfaces requires understanding how the human brain encodes stimuli across modalities such as visual, language (or text), etc. Brain encoding aims at constructing fMRI brain activity given a stimulus. There exists a plethora of neural encoding models which study brain encoding for single mode stimuli: visual (pretrained CNNs) or text (pretrained language models). Few recent papers have also obtained separate visual and text representation models and performed late-fusion using simple heuristics. However, previous work has failed to explore: (a) the effectiveness of image Transformer models for encoding visual stimuli, and (b) co-attentive multi-modal modeling for visual and text reasoning. In this paper, we systematically explore the efficacy of image Transformers (ViT, DEiT, and BEiT) and multi-modal Transformers (VisualBERT, LXMERT, and CLIP) for brain encoding. Extensive experiments on two popular datasets, BOLD5000 and Pereira, provide the following insights. (1) To the best of our knowledge, we are the first to investigate the effectiveness of image and multi-modal Transformers for brain encoding. (2) We find that VisualBERT, a multi-modal Transformer, significantly outperforms previously proposed single-mode CNNs, image Transformers as well as other previously proposed multi-modal models, thereby establishing new state-of-the-art. The supremacy of visio-linguistic models raises the question of whether the responses elicited in the visual regions are affected implicitly by linguistic processing even when passively viewing images. Future fMRI tasks can verify this computational insight in an appropriate experimental setting.
研究动机与目标
- 探究图像和多模态Transformer在从视觉和语言刺激编码fMRI脑活动方面的有效性。
- 确定联合视觉-语言建模是否能超越单模态方法,提升脑编码性能。
- 探索在被动观看图像时,语言处理是否仍会影响视觉脑区响应。
- 为多模态表征与人类脑响应之间的对齐提供计算见解。
提出的方法
- 本研究采用视觉-语言Transformer(VisualBERT、LXMERT、CLIP)和图像专用Transformer(ViT、DEiT、BEiT)来编码刺激。
- 通过跨注意力机制在多个层级上联合建模视觉和语言特征。
- 使用这些模型最后几层的表征来预测fMRI活动,不进行人工层选择。
- 实验在两个公开的fMRI数据集上进行:BOLD5000(纯视觉刺激)和Pereira(多模态刺激)。
- 通过计算预测与实际fMRI响应在脑区间的皮尔逊相关系数(PC)评估性能。
- 分析包括对具体概念与抽象概念的消融研究,以及在不同脑网络(如DMN、TP、视觉区域)中的比较。

实验结果
研究问题
- RQ1多模态Transformer(如VisualBERT)是否能在fMRI脑编码任务中超越单模态模型(CNN、图像Transformer)?
- RQ2即使仅呈现视觉刺激,语言处理的引入是否仍能提升脑活动预测性能?
- RQ3不同脑区(如视觉区、语言区、DMN)对视觉-语言表征的响应有何差异?
- RQ4在编码性能上,具体概念与抽象概念是否存在差异?
主要发现
- VisualBERT在BOLD5000和Pereira数据集上均取得最高性能,创下视觉-语言脑编码的新最先进水平。
- 如VisualBERT和LXMERT等多模态模型在fMRI预测任务中显著优于单模态模型(CNN和图像Transformer)。
- 视觉脑区(如Vision_Object、Vision_Face)与预测fMRI响应的相关性高于初级视觉区域。
- 在‘具体概念训练-抽象概念测试’设置下,PC得分优于‘抽象概念训练-具体概念测试’,表明从具体概念中学习更优。
- 场景选择性区域(RSC、OPA)在COCO-Scenes、ImageNet-Scenes和Scenes-Scenes任务中表现出更高相关性,尤其当模型在ImageNet或COCO上预训练时。
- 本研究预测,主动任务(如命名或决策)会引发比被动观看更强且更集中的视觉脑激活,提示语言可能在调节视觉处理中发挥作用。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。