[論文レビュー] Visio-Linguistic Brain Encoding
本論文は、視覚言語トランスフォーマー(例:VisualBERT、LXMERT)を用いて、視覚的および言語的刺激からのfMRI脳活動を予測する視覚言語的脳符号化を調査する。BOLD5000およびPereiraデータセットにおいて、マルチモodalトランスフォーマーが単一モodalモデル(CNN、画像トランスフォーマー)を著しく上回ることを示し、新たな最先端の結果を確立するとともに、被動的な画像視聴時でさえ言語処理が視覚的脳応答に間接的に影響を与える可能性があることを示唆している。
Enabling effective brain-computer interfaces requires understanding how the human brain encodes stimuli across modalities such as visual, language (or text), etc. Brain encoding aims at constructing fMRI brain activity given a stimulus. There exists a plethora of neural encoding models which study brain encoding for single mode stimuli: visual (pretrained CNNs) or text (pretrained language models). Few recent papers have also obtained separate visual and text representation models and performed late-fusion using simple heuristics. However, previous work has failed to explore: (a) the effectiveness of image Transformer models for encoding visual stimuli, and (b) co-attentive multi-modal modeling for visual and text reasoning. In this paper, we systematically explore the efficacy of image Transformers (ViT, DEiT, and BEiT) and multi-modal Transformers (VisualBERT, LXMERT, and CLIP) for brain encoding. Extensive experiments on two popular datasets, BOLD5000 and Pereira, provide the following insights. (1) To the best of our knowledge, we are the first to investigate the effectiveness of image and multi-modal Transformers for brain encoding. (2) We find that VisualBERT, a multi-modal Transformer, significantly outperforms previously proposed single-mode CNNs, image Transformers as well as other previously proposed multi-modal models, thereby establishing new state-of-the-art. The supremacy of visio-linguistic models raises the question of whether the responses elicited in the visual regions are affected implicitly by linguistic processing even when passively viewing images. Future fMRI tasks can verify this computational insight in an appropriate experimental setting.
研究の動機と目的
- 視覚的および言語的刺激からのfMRI脳活動の符号化において、画像およびマルチモダルトランスフォーマーの有効性を調査すること。
- 統合的視覚言語モデリングが、単一モダルアプローチを上回る脳符号化を実現するかどうかを特定すること。
- 被動的な画像視聴時でさえ、言語処理が視覚的脳領域に影響を与えるかどうかを調査すること。
- マルチモダル表現と人間の脳応答との整合性に関する計算的洞察を提供すること。
提案手法
- 本研究では、視覚言語トランスフォーマー(VisualBERT、LXMERT、CLIP)および画像専用トランスフォーマー(ViT、DEiT、BEiT)を用いて刺激を符号化する。
- 複数の層で視覚的および言語的特徴を同時にモデル化するために、クロスアテンションメカニズムを用いる。
- fMRI活動は、これらのモデルの最終層からの表現を用いて予測され、手動による層選択は一切行われない。
- 実験は、2つの公開fMRIデータセット(BOLD5000:純粋に視覚的刺激、Pereira:マルチモーダル刺激)を用いて実施される。
- 性能評価には、脳領域全体における予測されたfMRI応答と実際の応答の間のピアソン相関(PC)スコアが用いられる。
- 分析には、具体的な概念と抽象的な概念のアブレーションスタディおよび脳ネットワーク(例:DMN、TP、視覚領域)間の比較が含まれる。

実験結果
リサーチクエスチョン
- RQ1VisualBERTのようなマルチモーダルトランスフォーマーは、CNN や画像トランスフォーマーといった単一モダルモデルを上回ってfMRI脳符号化を達成できるか?
- RQ2視覚的刺激のみが提示された場合でも、言語処理の統合が脳活動予測を向上させるか?
- RQ3視覚的、言語的、DMNなどの異なる脳領域は、視覚言語的表現に対してどのように反応するか?
- RQ4具体的な概念と抽象的な概念の間で、符号化性能に差異があるか?
主な発見
- VisualBERTは、BOLD5000およびPereiraデータセットの両方で最高の性能を示し、視覚言語的脳符号化分野における新たな最先端の結果を樹立した。
- VisualBERT や LXMERT などのマルチモーダルモデルは、CNN や画像トランスフォーマーといった単一モダルモデルを著しく上回る性能を示した。
- 視覚的脳領域(例:Vision_Object、Vision_Face)は、初期視覚領域よりも予測fMRI応答との相関が高かった。
- 具体的な概念で学習し、抽象的な概念でテストするモデル(concrete-train-abstract-test)は、逆の順序(abstract-train-concrete-test)よりも高いPCスコアを示し、具体的な概念からの学習がより効果的であることを示した。
- シーン選択的領域(RSC、OPA)は、COCO-Scenes、ImageNet-Scenes、Scenes-Scenes タスクにおいて、特にImageNet や COCO で事前学習されたモデルで高い相関を示した。
- 本研究では、アクティブタスク(例:名前をつける、意思決定を行う)は、被動的視聴よりも強力で集中した視覚的脳活性化を引き起こすと予測しており、言語が視覚処理を調整する役割を果たしている可能性を示唆している。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。