Skip to main content
QUICK REVIEW

[논문 리뷰] Cinematic Mindscapes: High-quality Video Reconstruction from Brain Activity

Zijiao Chen, Jiaxin Qing|arXiv (Cornell University)|2023. 05. 19.
Functional Brain Connectivity Studies인용 수 20
한 줄 요약

MinD-Video는 fMRI를 통해 고품질의 의미적으로 의미 있는 비디오를 재구성합니다. fMRI 인코더를 보강된 안정적 확산 비디오 생성기에 분리하는 방식으로, 점진적이며 다중 모달적이고 적대적 지침을 사용합니다. 이는 최첨단 의미 정확도와 경쟁력 있는 픽셀 충실도를 달성하며 어텐션 맵을 통한 해석가능성을 제공합니다.

ABSTRACT

Reconstructing human vision from brain activities has been an appealing task that helps to understand our cognitive process. Even though recent research has seen great success in reconstructing static images from non-invasive brain recordings, work on recovering continuous visual experiences in the form of videos is limited. In this work, we propose Mind-Video that learns spatiotemporal information from continuous fMRI data of the cerebral cortex progressively through masked brain modeling, multimodal contrastive learning with spatiotemporal attention, and co-training with an augmented Stable Diffusion model that incorporates network temporal inflation. We show that high-quality videos of arbitrary frame rates can be reconstructed with Mind-Video using adversarial guidance. The recovered videos were evaluated with various semantic and pixel-level metrics. We achieved an average accuracy of 85% in semantic classification tasks and 0.19 in structural similarity index (SSIM), outperforming the previous state-of-the-art by 45%. We also show that our model is biologically plausible and interpretable, reflecting established physiological processes.

연구 동기 및 목표

  • 비침습적 뇌활동(fMRI)으로부터 연속적인 시각 경험(비디오)을 재구성하는 방법을 이해한다.
  • 품질과 유연성을 향상시키기 위해 fMRI 인코딩을 비디오 생성으로부터 분리한 두 모듈 파이프라인을 개발한다.
  • fMRI의 시간 해상도 차이를 해소하기 위해 점진적이고 다중 모달 학습과 시간적 어텐션을 활용한다.
  • 씬 동적 어텐션(scene-dynamic attention)과 적대적 지침으로 Stable Diffusion 기반 비디오 생성기를 보강하여 충실도를 향상시킨다.

제안 방법

  • 두 모듈 파이프라인: 보강된 Stable Diffusion 비디오 생성기에서 fMRI 인코더를 별도로 학습한 다음 두 모듈을 공동 학습한다.
  • 점진적 학습: 대규모 MBM 사전 학습에 이어, 윈도우화된 fMRI에서 시공간 어텐션을 활용한 다중 모달 대조 학습.
  • 혈역학적 지연을 고려하고 슬라이딩 윈도우 fMRI를 처리하기 위한 시공간 어텐션.
  • 씬-다이나믹 희소 CA(SC) 어텐션을 갖춘 보강된 Stable Diffusion으로 이전 프레임에 조건을 걸면서 장면 변화도 허용.
  • 부정 조건부를 통한 적대적 지침으로 fMRI 조건부 샘플링 품질을 향상.
  • 해석 가능성을 바탕으로 뇌 데이터로부터 학습: 어텐션 맵을 시각화하여 디코더 전략을 뇌 네트워크에 매핑한다.
Figure 1 : Brain decoding & video reconstruction . We propose a progressive learning approach to recover continuous visual experience from fMRI. High-quality videos with accurate semantics and motions are reconstructed.
Figure 1 : Brain decoding & video reconstruction . We propose a progressive learning approach to recover continuous visual experience from fMRI. High-quality videos with accurate semantics and motions are reconstructed.

실험 결과

연구 질문

  • RQ1HRF 지연에도 불구하고 임의의 프레임 레이트에서 연속 비디오 콘텐츠를 fMRI로부터 재구성할 수 있는가?
  • RQ2점진적이고 다중 모달 fMRI 인코딩을 공동 학습된 비디오 생성기와 결합하면, 이전 방법에 비해 의미적 및 픽셀 수준 충실도가 향상되는가?
  • RQ3적대적 지침이 조건화 효과와 생성된 비디오의 다양성에 어떤 영향을 미치는가?
  • RQ4어텐션 맵이 해독된 시각 콘텐츠에 기여하는 뇌 영역과 네트워크에 대해 무엇을 보여주는가?

주요 결과

  • 본 방법은 비디오 콘텐츠에서 85%의 의미 분류 정확도와 0.19의 SSIM을 달성하며, 이전의 최첨단 대비 45% 향상됐다.
  • 본 방법은 피실험자 간의 동적 움직임과 장면 역학이 정확한 고품질 비디오를 생성한다.
  • 어텐션 분석은 시각 피질의 지배를 보이며 상위 인지 네트워크의 기여와 함께 생물학적 가능성과 일치한다.
  • 점진적 학습 단계는 국부적 특징에서 전역적 특징으로의 전이를 반영하며, 후반 계층은 추상적 의미 정보에 집중한다.
  • Ablation 연구는 성능을 위해 윈도우 크기, 다중 모달 대조 학습, 및 적대적 지침의 중요성을 보여준다.
  • 본 프레임워크는 장면 전환을 포함한 다양한 장면과 모션을 재구성하되 프레임 일관성을 유지한다.
Figure 2 : MinD-Video Overview . Our method has two modules that are trained separately, then finetuned together. The fMRI encoder progressively learns fMRI features through multiple stages, including SC-MBM pre-training and multimodal contrastive learning. A spatiotemporal attention is designed to
Figure 2 : MinD-Video Overview . Our method has two modules that are trained separately, then finetuned together. The fMRI encoder progressively learns fMRI features through multiple stages, including SC-MBM pre-training and multimodal contrastive learning. A spatiotemporal attention is designed to

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.