[논문 리뷰] Deep Neural Networks and Brain Alignment: Brain Encoding and Decoding (Survey)
이 종합 검토는 fMRI 데이터를 사용한 딥 뉴럴 네트워크(DNN) 기반 뇌 인코딩 및 디코딩 모델을 다루며, DNN이 자극(텍스트, 이미지, 음성)에서 의미적 표현을 어떻게 학습하고 뇌 활동으로 매핑하는지에 중점을 둡니다. 최근의 DNN 아키텍처, 자극 표현 방식, 데이터셋에 대한 발전을 종합적으로 분석하며, 의미 벡터 복원 및 자극 재구성의 향상 사례를 강조합니다. 뇌-컴퓨터 인터페이스 및 인지신경과학 분야의 적용 가능성을 제시합니다.
Can artificial intelligence unlock the secrets of the human brain? How do the inner mechanisms of deep learning models relate to our neural circuits? Is it possible to enhance AI by tapping into the power of brain recordings? These captivating questions lie at the heart of an emerging field at the intersection of neuroscience and artificial intelligence. Our survey dives into this exciting domain, focusing on human brain recording studies and cutting-edge cognitive neuroscience datasets that capture brain activity during natural language processing, visual perception, and auditory experiences. We explore two fundamental approaches: encoding models, which attempt to generate brain activity patterns from sensory inputs; and decoding models, which aim to reconstruct our thoughts and perceptions from neural signals. These techniques not only promise breakthroughs in neurological diagnostics and brain-computer interfaces but also offer a window into the very nature of cognition. In this survey, we first discuss popular representations of language, vision, and speech stimuli, and present a summary of neuroscience datasets. We then review how the recent advances in deep learning transformed this field, by investigating the popular deep learning based encoding and decoding architectures, noting their benefits and limitations across different sensory modalities. From text to images, speech to videos, we investigate how these models capture the brain's response to our complex, multimodal world. While our primary focus is on human studies, we also highlight the crucial role of animal models in advancing our understanding of neural mechanisms. Throughout, we mention the ethical implications of these powerful technologies, addressing concerns about privacy and cognitive liberty. We conclude with a summary and discussion of future trends in this rapidly evolving field.
연구 동기 및 목표
- 다양한 모odalities(텍스트, 시각, 언어)에 걸쳐 심층학습 기반 뇌 인코딩 및 디코딩 분야의 최근 발전을 종합하기 위해.
- 자연주의 자극에서 뇌 표현을 모델링하는 데 있어 심층신경망 아키텍처의 역할을 분석하기 위해.
- fMRI 활동을 예측하는 데 있어 다양한 자극 표현 방식(예: BERT, 워드 임베딩, 비전 트랜스포머)의 효과성을 평가하기 위해.
- 주로 수동 자극 처리 및 L2 언어 이해와 관련된 문제점을 포함해 현재 자연주의 fMRI 연구의 한계를 특정하기 위해.
- 뇌 기반 인공지능, 다중모달 디코딩, 신경적 내성 있는 신경망 설계 분야의 향후 연구 방향을 제시하기 위해.
제안 방법
- fMRI를 활용한 뇌 인코딩 및 디코딩에 관한 230개 이상의 연구를 대상으로 한 체계적 검토.
- 모달리티(텍스트, 이미지, 음성, 영상) 및 자극 유형(내용물, 영화, 자연주의 자극)에 따라 신경과학 데이터셋을 분류.
- 자극 표현 기법 분석: 분산 워드 임베딩, 문장 트랜스포머(예: InferSent, BERT), 비전 트랜스포머.
- 회귀 함수 e: R → F를 사용한 인코딩 모델 평가: 의미 표현에서 fMRI 활동을 예측.
- 뇌 활동에서 의미 표현을 재구성하는 데 사용되는 함수 d: F → R을 활용한 디코딩 모델 평가. 다중시점 및 크로스시점 디코딩 설정 포함.
- 성능 향상을 위해 최근 아키텍처인 트랜스포머 및 피니튜닝 전략(예: 뇌 디코딩을 위한 BERT 피니튜닝)을 통합.

실험 결과
연구 질문
- RQ1전통적 모델 대비 심층신경망이 뇌 인코딩 및 디코딩 정확도를 어떻게 향상시키는가?
- RQ2어떤 종류의 자극 표현 방식(예: BERT, 워드 임베딩, 비전 트랜스포머)이 fMRI 뇌 활동과 가장 잘 일치하는가?
- RQ3뇌 디코딩 모델이 fMRI 데이터에서 의미 벡터나 심지어 전체 자극(예: 이미지, 문장)을 얼마나 잘 재구성할 수 있는가?
- RQ4다중시점 및 크로스시점 디코딩 프레임워크는 뇌 디코딩 모델의 일반화 능력과 내성 강화에 어떻게 기여하는가?
- RQ5수동 자극 노출 중 진정한 인지 처리를 반영하지 못하는 현재 자연주의 fMRI 프로토콜의 주요 한계는 무엇인가?
주요 결과
- BERT 및 InferSent와 같은 트랜스포머 기반 모델은 자연어 이해(NLU) 작업에서 피니튜닝을 거치면 뇌 디코딩 성능이 크게 향상됨.
- 크로스시점 디코딩(CVD)은 한 모달리티(예: 텍스트)의 의미 표현을 다른 모달리티(예: 이미지) 처리 중 기록된 뇌 활동에서 디코딩할 수 있게 하며, 이미지 캡션 생성 및 키워드 추출과 같은 작업을 가능하게 함.
- 다중시점 디코딩(MVD) 모델은 다양한 자극 시점 간의 일반화 능력을 향상시켜 입력 모달리티 변동에 대한 내성 강화에 기여함.
- 구문을 최소화한 의미 중심 표현을 생성하는 모델일수록 뇌 활동과 더 잘 일치함을 시사하며, 더 단순하고 의미 중심의 표현 방식이 뇌 활동과 더 잘 부합함.
- 최근 모델들은 의미 벡터 외에도 연속적인 언어, 이미지, 심지어 음성까지 fMRI 데이터에서 재구성할 수 있으며, 이는 다중모달 디코딩의 실현 가능성을 입증함.
- 수동 청취 중 뇌 활동을 해석하는 데는 여전히 한계가 있으며, 특히 L2 언어 처리에서는 뇌 활동이 L1 억제를 반영할 수 있으므로 L2 이해를 반영하지 못할 수 있음.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.