Skip to main content
QUICK REVIEW

[论文解读] Multimodal foundation models are better simulators of the human brain

Haoyu Lu, Qiongyi Zhou|arXiv (Cornell University)|Aug 17, 2022
Domain Adaptation and Few-Shot Learning被引用 11
一句话总结

本文提出,与单模态模型相比,多模态基础模型(如BriVL)作为人类大脑的模拟器更为优越。通过在1500万对图像-文本数据上训练大规模多模态模型,并将其表征与fMRI数据进行对比,研究发现,经过多模态训练的视觉和语言编码器在与多感官整合相关的脑区(特别是腹侧颞叶皮层和后中颞回)中,能更准确地预测神经响应。

ABSTRACT

Multimodal learning, especially large-scale multimodal pre-training, has developed rapidly over the past few years and led to the greatest advances in artificial intelligence (AI). Despite its effectiveness, understanding the underlying mechanism of multimodal pre-training models still remains a grand challenge. Revealing the explainability of such models is likely to enable breakthroughs of novel learning paradigms in the AI field. To this end, given the multimodal nature of the human brain, we propose to explore the explainability of multimodal learning models with the aid of non-invasive brain imaging technologies such as functional magnetic resonance imaging (fMRI). Concretely, we first present a newly-designed multimodal foundation model pre-trained on 15 million image-text pairs, which has shown strong multimodal understanding and generalization abilities in a variety of cognitive downstream tasks. Further, from the perspective of neural encoding (based on our foundation model), we find that both visual and lingual encoders trained multimodally are more brain-like compared with unimodal ones. Particularly, we identify a number of brain regions where multimodally-trained encoders demonstrate better neural encoding performance. This is consistent with the findings in existing studies on exploring brain multi-sensory integration. Therefore, we believe that multimodal foundation models are more suitable tools for neuroscientists to study the multimodal signal processing mechanisms in the human brain. Our findings also demonstrate the potential of multimodal foundation models as ideal computational simulators to promote both AI-for-brain and brain-for-AI research.

研究动机与目标

  • 探究多模态基础模型是否比单模态模型更优地模拟人类大脑反应。
  • 评估多模态训练编码器(视觉与语言)在人类受试者fMRI数据中的神经编码性能。
  • 识别多模态预训练导致表征更具脑似性的特定脑区。
  • 探索多模态基础模型作为脑启发式人工智能与神经科学研究计算模拟器的潜力。

提出的方法

  • 使用对比学习在1500万对图像-文本对上预训练大规模多模态基础模型BriVL。
  • 从BriVL的编码器中提取视觉与语言特征,并与单模态训练模型(ViT与BERT)的特征进行比较。
  • 应用带状岭回归模型来建模fMRI数据中的神经响应,以模型特征作为预测因子。
  • 使用决定系数(R²)量化模型特征在不同体素中解释神经响应方差的程度。
  • 对R²进行分解,以分离视觉与语言特征对神经编码性能的独立贡献。
  • 在多个脑区评估编码性能,重点关注与多感官整合相关的区域。

实验结果

研究问题

  • RQ1多模态基础模型是否产生比单模态模型更优的、能更好预测人类大脑活动的表征?
  • RQ2在哪些脑区中,使用多模态训练编码器时,神经编码性能显著提升?
  • RQ3与单模态模型相比,视觉与语言特征对神经编码的贡献在多模态与单模态模型中存在何种差异?
  • RQ4多模态表征在多大程度上与人类大脑中已知的多感官整合模式相吻合?
  • RQ5多模态基础模型能否作为人工智能与神经科学研究的有效计算模拟器?

主要发现

  • 与单模态训练的ViT和BERT相比,BriVL的多模态训练视觉与语言编码器在预测fMRI响应时表现出显著更高的R²值。
  • BriVL的视觉编码器在腹侧颞叶皮层和后中颞回表现出更优的神经编码性能,这些区域与多感官整合相关。
  • BriVL的语言编码器在左后中颞回表现出更强的编码能力,该区域与语义处理相关。
  • 在多个脑区中,BriVL中视觉与语言特征的联合贡献,解释了比单模态模型更多的神经响应独特方差。
  • 带状岭回归分析表明,多模态预训练增强了模型特征与人类神经响应表征对齐的能力。
  • 本研究证实,多模态基础模型作为模拟器更具脑似性,尤其在涉及跨模态整合的脑区中表现更优。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。