Skip to main content
QUICK REVIEW

[论文解读] Deep Neural Networks and Brain Alignment: Brain Encoding and Decoding (Survey)

Subba Reddy Oota, Chen, Zijiao|arXiv (Cornell University)|Jul 17, 2023
EEG and Brain-Computer Interfaces参考文献 202被引用 5
一句话总结

本综述回顾了基于功能性磁共振成像(fMRI)数据的深度神经网络(DNN)脑编码与解码模型,重点探讨DNN如何从刺激(文本、图像、音频)中学习语义表征,并将其映射到脑活动。该综述整合了DNN架构、刺激表征和数据集的最新进展,突出展示了在解码语义向量和重建刺激方面的改进,其应用涵盖脑机接口与认知神经科学。

ABSTRACT

Can artificial intelligence unlock the secrets of the human brain? How do the inner mechanisms of deep learning models relate to our neural circuits? Is it possible to enhance AI by tapping into the power of brain recordings? These captivating questions lie at the heart of an emerging field at the intersection of neuroscience and artificial intelligence. Our survey dives into this exciting domain, focusing on human brain recording studies and cutting-edge cognitive neuroscience datasets that capture brain activity during natural language processing, visual perception, and auditory experiences. We explore two fundamental approaches: encoding models, which attempt to generate brain activity patterns from sensory inputs; and decoding models, which aim to reconstruct our thoughts and perceptions from neural signals. These techniques not only promise breakthroughs in neurological diagnostics and brain-computer interfaces but also offer a window into the very nature of cognition. In this survey, we first discuss popular representations of language, vision, and speech stimuli, and present a summary of neuroscience datasets. We then review how the recent advances in deep learning transformed this field, by investigating the popular deep learning based encoding and decoding architectures, noting their benefits and limitations across different sensory modalities. From text to images, speech to videos, we investigate how these models capture the brain's response to our complex, multimodal world. While our primary focus is on human studies, we also highlight the crucial role of animal models in advancing our understanding of neural mechanisms. Throughout, we mention the ethical implications of these powerful technologies, addressing concerns about privacy and cognitive liberty. We conclude with a summary and discussion of future trends in this rapidly evolving field.

研究动机与目标

  • 整合多模态(文本、视觉、语音)深度学习驱动的脑编码与解码的最新进展。
  • 分析深度神经网络架构在建模自然刺激下脑表征中的作用。
  • 评估不同刺激表征方式(如BERT、词嵌入、视觉Transformer)在预测fMRI活动方面的有效性。
  • 识别当前自然刺激fMRI研究的局限性,特别是被动刺激处理与第二语言理解方面的不足。
  • 提出脑启发式人工智能、多模态解码与神经鲁棒神经网络设计的未来研究方向。

提出的方法

  • 对230余项基于fMRI的脑编码与解码研究进行系统性综述,重点关注深度学习模型。
  • 按模态(文本、图像、音频、视频)和刺激类型(叙事、电影、自然刺激)对神经科学数据集进行分类。
  • 分析刺激表征技术:分布式词嵌入、句子编码器(如InferSent、BERT)以及视觉Transformer。
  • 使用回归函数 e: R → F 评估编码模型,从语义表征预测fMRI活动。
  • 使用函数 d: F → R 评估解码模型,从脑活动重建语义表征,包括多视角与跨视角解码设置。
  • 引入近期架构如Transformer及微调策略(如BERT在脑解码中的微调)以提升性能。
Figure 1: Computational Cognitive Neuroscience of Brain Encoding and Decoding: Datasets & Stimulus Representations
Figure 1: Computational Cognitive Neuroscience of Brain Encoding and Decoding: Datasets & Stimulus Representations

实验结果

研究问题

  • RQ1与传统模型相比,深度神经网络在提升脑编码与解码准确性方面有何改进?
  • RQ2何种类型的刺激表征(如BERT、词嵌入、视觉Transformer)能与fMRI脑活动实现最佳对齐?
  • RQ3脑解码模型在多大程度上能从fMRI数据中重建语义向量,甚至完整刺激(如图像、句子)?
  • RQ4多视角与跨视角解码框架如何增强脑解码模型的泛化能力与鲁棒性?
  • RQ5当前自然刺激fMRI范式在捕捉被动刺激暴露期间真实认知加工方面存在哪些关键局限?

主要发现

  • 基于Transformer的模型(如BERT与InferSent)在经过自然语言理解任务微调后,显著提升了脑解码性能。
  • 跨视角解码(CVD)可实现从一种模态(如图像)处理过程中记录的脑活动,解码出另一种模态(如文本)的语义表征,支持图像字幕生成与关键词提取等任务。
  • 多视角解码(MVD)模型可在不同刺激视角间实现泛化,提升对输入模态变化的鲁棒性。
  • 当模型生成语法信息较少的语义表征时,解码性能更高,表明更简洁、更侧重语义的表征与脑活动更匹配。
  • 近期模型不仅能重建语义向量,还能从fMRI中恢复连续语言、图像甚至语音,证明了多模态解码的可行性。
  • 在被动听觉刺激处理中,尤其是第二语言理解时,脑活动的解释仍存在局限,脑活动可能反映第一语言抑制而非第二语言理解。
Figure 2: Representative Samples of Naturalistic Brain Dataset: (LEFT) Brain activity recorded when subjects are reading and listening to the same narrative ( Deniz et al. 2019 ), and (RIGHT) example naturalistic image stimuli from various public repositories: BOLD5000 ( Chang et al. 2019 ), SSfMRI
Figure 2: Representative Samples of Naturalistic Brain Dataset: (LEFT) Brain activity recorded when subjects are reading and listening to the same narrative ( Deniz et al. 2019 ), and (RIGHT) example naturalistic image stimuli from various public repositories: BOLD5000 ( Chang et al. 2019 ), SSfMRI

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。