[논문 리뷰] Contrast, Attend and Diffuse to Decode High-Resolution Images from Brain Activities
본 논문은 두 단계로 구성된 fMRI 표현 학습 프레임워크(DC-MAE 사전 학습 및 교차 모달리티 튜닝)로 잠재 확산 모델을 조건화하여 뇌 활동으로부터 고해상도 영상 재구성을 수행하고, 기존 방법들에 비해 상당한 성과를 보인다.
Decoding visual stimuli from neural responses recorded by functional Magnetic Resonance Imaging (fMRI) presents an intriguing intersection between cognitive neuroscience and machine learning, promising advancements in understanding human visual perception and building non-invasive brain-machine interfaces. However, the task is challenging due to the noisy nature of fMRI signals and the intricate pattern of brain visual representations. To mitigate these challenges, we introduce a two-phase fMRI representation learning framework. The first phase pre-trains an fMRI feature learner with a proposed Double-contrastive Mask Auto-encoder to learn denoised representations. The second phase tunes the feature learner to attend to neural activation patterns most informative for visual reconstruction with guidance from an image auto-encoder. The optimized fMRI feature learner then conditions a latent diffusion model to reconstruct image stimuli from brain activities. Experimental results demonstrate our model's superiority in generating high-resolution and semantically accurate images, substantially exceeding previous state-of-the-art methods by 39.34% in the 50-way-top-1 semantic classification accuracy. Our research invites further exploration of the decoding task's potential and contributes to the development of non-invasive brain-machine interfaces.
연구 동기 및 목표
- 뇌 fMRI 표현의 노이즈 제거 및 정제를 통해 디코딩 품질을 향상시킨다.
- 시각 관련 뇌 신호에 주의하도록 교차 모달리티 가이던스를 활용한다.
- 뇌 활동으로부터 고해상도 의미적으로 충실한 이미지를 생성하기 위해 잠재 확산 모델을 통합한다.
- 기준 방법들에 비해 GOD 및 BOLD5000 데이터셋에서 재구성 성능이 우수함을 입증한다.
제안 방법
- 라벨링되지 않은 fMRI 데이터에서 Double-contrastive Masked Auto-Encoder (DC-MAE)로 fMRI 특징 학습기를 사전 학습하여 노이즈 제거된 표현을 학습한다.
- Phase 2에서 재구성을 위한 정보가 풍부한 뇌 패턴에 주의를 유도하기 위해 교차 어텐션을 사용하는 이미지 자동 인코더로 fMRI 인코더를 미세 조정한다.
- 최적화된 fMRI 인코더로 잠재 확산 모델(LDM)을 조건시켜 뇌 활동으로부터 이미지를 생성한다.
- fMRI 특징을 이용한 교차 어텐션 컨디셔닝으로 LDM을 미세 조정하여 조건부 이미지 생성을 수행한다.
- 사전 학습된 ImageNet 분류기를 통해 50-way-top-1 의미 정확도에 기반한 평가 지표를 사용하여 의미적 정확성을 평가한다.

실험 결과
연구 질문
- RQ1DC-MAE가 개인 간 fMRI 표현의 노이즈를 효과적으로 제거하고 정렬할 수 있는가?
- RQ2fMRI와 이미지 자동 인코더 간의 교차 모달리티 지도가 fMRI에서 시각적 재구성의 품질을 향상시키는가?
- RQ3fMRI로 조건화된 잠재 확산 모델이 고해상도이고 의미적으로 정확한 이미지를 생성하는가?
- RQ4FRL Phase 2에서 재구성 손실 및 마스크 비율이 디코딩 성능에 미치는 상대적 기여는 무엇인가?
- RQ5제안된 프레임워크가 GOD 및 BOLD5000와 같은 데이터셋에서 확장 가능하는가?
주요 결과
- 제안된 모델은 GOD 및 BOLD5000 데이터에서 50-way-top-1 정확도 면에서 이전 최첨단 방법보다 39.34%를 상회한다.
- 두 단계 FRL(DC-MAE 사전 학습에 이어 교차 모달리티 튜닝)이 고해상도이며 의미적으로 정확한 재구성을 제공한다.
- 결합 fMRI 및 이미지 재구성 손실과 신중하게 선택된 마스크 비율 및 디코더 깊이가 성능에 결정적임을 보여주는 제거 연구들.
- 방법은 GOD 피험자 CSI1에서 50-way-top-1 정확도 25를 달성하고, 테스트 세트 fMRI 데이터를 튜닝에 사용하지 않았을 때 DC-LDM에 비해 GOD 피험자 1,2,4,5에서 우수한 성능을 보인다.
- 이 접근법은 LDM 학습 시 데이터셋으로 인한 편향과 디테일 재구성 간의 균형 문제를 식별한다.
![Figure 2: [a] Demo of the forward and backward processes of the diffusion model. [b] The forward process of the diffusion model which progressively corrupts an image with Gaussian noise. [c] In the backward process, the diffusion model, conditioned on our pretrained fMRI encoder, gradually denoises](https://ar5iv.labs.arxiv.org/html/2305.17214/assets/x2.png)
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.