[论文解读] More Than Meets the Eye: Self-Supervised Depth Reconstruction From Brain Activity
本文提出了一种首次直接从fMRI脑活动重建密集3D深度图的方法,采用自监督学习。通过利用未配对自然图像上的预训练单目深度估计模型,并引入基于深度的感知相似性度量,该方法在直接深度预测方面优于间接从重建图像中恢复深度的方法,在1000选1识别任务中,直接深度预测的平均深度排名达到97,显著优于间接方法。
In the past few years, significant advancements were made in reconstruction of observed natural images from fMRI brain recordings using deep-learning tools. Here, for the first time, we show that dense 3D depth maps of observed 2D natural images can also be recovered directly from fMRI brain recordings. We use an off-the-shelf method to estimate the unknown depth maps of natural images. This is applied to both: (i) the small number of images presented to subjects in an fMRI scanner (images for which we have fMRI recordings - referred to as "paired" data), and (ii) a very large number of natural images with no fMRI recordings ("unpaired data"). The estimated depth maps are then used as an auxiliary reconstruction criterion to train for depth reconstruction directly from fMRI. We propose two main approaches: Depth-only recovery and joint image-depth RGBD recovery. Because the number of available "paired" training data (images with fMRI) is small, we enrich the training data via self-supervised cycle-consistent training on many "unpaired" data (natural images & depth maps without fMRI). This is achieved using our newly defined and trained Depth-based Perceptual Similarity metric as a reconstruction criterion. We show that predicting the depth map directly from fMRI outperforms its indirect sequential recovery from the reconstructed images. We further show that activations from early cortical visual areas dominate our depth reconstruction results, and propose means to characterize fMRI voxels by their degree of depth-information tuning. This work adds an important layer of decoded information, extending the current envelope of visual brain decoding capabilities.
研究动机与目标
- 将基于fMRI的视觉解码从RGB图像重建扩展到密集3D深度图的重建。
- 通过利用大规模未配对自然图像及其估计深度图,解决fMRI-图像配对数据稀缺的问题。
- 开发一种自监督训练框架,通过感知一致性提升从fMRI中重建深度的能力。
- 通过其对深度信息与颜色/RGB内容的敏感性,对fMRI体素进行表征。
提出的方法
- 使用预训练的单目深度模型(MiDaS)为配对fMRI图像和大量未配对自然图像估计深度图。
- 使用未配对数据上的循环一致性自监督方法,训练一个深度编码器-解码器网络,将fMRI活动映射为深度图(或RGBD)。
- 提出一种新颖的基于深度的感知相似性度量,利用预训练深度网络的特征来引导重建保真度。
- 使用基于感知的损失函数,确保重建的深度图在感知上合理且与人类深度感知一致。
- 训练两种主要架构:仅深度恢复和联合RGBD恢复,共享编码器和解码器组件。
- 提出体素深度敏感性指数(VDSI),用于量化单个体素对深度信息与RGB信息编码的选择性。
实验结果
研究问题
- RQ1能否直接从fMRI脑活动重建密集3D深度图,而无需依赖中间图像重建?
- RQ2在未配对自然图像及其估计深度图上进行自监督训练,是否能提升基于fMRI的深度重建性能?
- RQ3与通过重建RGB图像间接恢复深度相比,直接从fMRI预测深度在准确性上表现如何?
- RQ4哪些脑区和体素对fMRI活动中的深度信息最敏感?
- RQ5能否系统性地表征单一个体素对深度信息的调谐程度?
主要发现
- 直接RGBD方法在1000选1识别任务中达到平均深度排名97,显著优于间接方法(分别为38和56的排名水平)。
- 深度重建性能主要由初级视觉皮层(LVC,V1–V3)体素驱动,而高级视觉皮层(HVC)体素对深度解码的贡献较小。
- 所提出的基于深度的感知相似性度量能有效引导重建,生成感知上合理的深度图。
- 体素深度敏感性指数(VDSI)在不同数据集中表现出强一致性,大多数体素对深度的敏感性远低于对颜色的敏感性。
- 直接深度预测方法在相同1000选1任务中实现了平均RGB排名16的高质量深度图,表明其在图像与深度联合重建方面表现优异。
- 本研究证明,深度是一种可行且信息丰富的模态,可用于基于fMRI的脑解码,拓展了从脑活动可恢复的视觉信息范围。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。