Skip to main content
QUICK REVIEW

[論文レビュー] Multimodal foundation models are better simulators of the human brain

Haoyu Lu, Qiongyi Zhou|arXiv (Cornell University)|Aug 17, 2022
Domain Adaptation and Few-Shot Learning被引用数 11
ひとこと要約

本論文は、BriVL などのマルチモーダル基盤モデルが、単モーダルモデルよりも人間の脳を優れたシミュレータとして機能することを提案している。1,500万枚の画像・テキストペアを用いて大規模なマルチモーダルモデルを訓練し、fMRIデータと照合して表現を評価することで、マルチモーダルに訓練された視覚的および言語的エンコーダーが、特に腹側側頭皮質および後側中側頭回といった、マルチセンサリ統合に関連する脳領域における神経応答をよりよく予測することが示された。

ABSTRACT

Multimodal learning, especially large-scale multimodal pre-training, has developed rapidly over the past few years and led to the greatest advances in artificial intelligence (AI). Despite its effectiveness, understanding the underlying mechanism of multimodal pre-training models still remains a grand challenge. Revealing the explainability of such models is likely to enable breakthroughs of novel learning paradigms in the AI field. To this end, given the multimodal nature of the human brain, we propose to explore the explainability of multimodal learning models with the aid of non-invasive brain imaging technologies such as functional magnetic resonance imaging (fMRI). Concretely, we first present a newly-designed multimodal foundation model pre-trained on 15 million image-text pairs, which has shown strong multimodal understanding and generalization abilities in a variety of cognitive downstream tasks. Further, from the perspective of neural encoding (based on our foundation model), we find that both visual and lingual encoders trained multimodally are more brain-like compared with unimodal ones. Particularly, we identify a number of brain regions where multimodally-trained encoders demonstrate better neural encoding performance. This is consistent with the findings in existing studies on exploring brain multi-sensory integration. Therefore, we believe that multimodal foundation models are more suitable tools for neuroscientists to study the multimodal signal processing mechanisms in the human brain. Our findings also demonstrate the potential of multimodal foundation models as ideal computational simulators to promote both AI-for-brain and brain-for-AI research.

研究の動機と目的

  • マルチモーダル基盤モデルが単モーダルモデルよりも人間の脳応答をよりよくシミュレートするかどうかを調査すること。
  • 人間被験者からのfMRIデータと照合して、マルチモーダルに訓練されたエンコーダー(視覚的および言語的)の神経符号化性能を評価すること。
  • マルチモーダル事前学習がより脳に類似した表現をもたらす特定の脳領域を同定すること。
  • マルチモーダル基盤モデルが、脳にインspiredされたAIおよび神経科学研究のための計算的シミュレータとしての可能性を検討すること。

提案手法

  • 対照的学習を用いて、1,500万枚の画像・テキストペア上で大規模なマルチモーダル基盤モデルBriVLを事前学習した。
  • BriVLのエンコーダーから視覚的および言語的特徴を抽出し、単モーダルに学習されたモデル(ViTおよびBERT)の特徴と比較した。
  • fMRIデータにおける神経応答をモデル化するためにバンド付きリッジ回帰を適用した。
  • モデル特徴を予測変数として用い、各ボクセルにおける神経応答の分散をどれだけうまく説明できるかを係数決定係数(R²)で定量化した。
  • R²を分割することで、視覚的および言語的特徴が神経符号化性能に果たす独自の寄与を分離した。
  • マルチセンサリ統合に関連する領域が特に注目されるように、複数の脳領域における符号化性能を評価した。

実験結果

リサーチクエスチョン

  • RQ1マルチモーダル基盤モデルは、単モーダルモデルよりも人間の脳活動をよりよく予測する表現を生成するか?
  • RQ2マルチモーダルに訓練されたエンコーダーを用いることで、どの脳領域で神経符号化性能が顕著に向上するか?
  • RQ3マルチモーダルモデルと単モーダルモデルの間で、視覚的および言語的特徴が神経符号化に与える寄与はどのように異なるか?
  • RQ4マルチモーダル表現は、人間脳におけるマルチセンサリ統合の既知のパターンとどの程度一致するか?
  • RQ5マルチモーダル基盤モデルは、AIおよび神経科学研究の両方の計算的シミュレータとして効果的であると見なせるか?

主な発見

  • BriVL から得られたマルチモーダルに訓練された視覚的および言語的エンコーダーは、単モーダルに訓練されたViTおよびBERTと比較して、fMRI応答を予測する際、有意に高いR²値を達成した。
  • BriVLの視覚的エンコーダーは、マルチセンサリ統合に関連する領域として知られる腹側側頭皮質および後側中側頭回で、優れた神経符号化性能を示した。
  • BriVLの言語的エンコーダーは、意味処理に関連する左側後側中側頭回で、より強い符号化を示した。
  • 複数の脳領域において、BriVLにおける視覚的および言語的特徴の組み合わせによる寄与が、単モーダルモデルよりも神経応答の独自分散をより多く説明した。
  • バンド付きリッジ回帰分析により、マルチモーダル事前学習がモデル特徴と人間の神経応答との間の表現一致を強化することが明らかになった。
  • 本研究は、マルチモーダル基盤モデルが、特にクロスモーダル統合に関与する領域において、より脳に類似したシミュレータであることを確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。