[Paper Review] More Than Meets the Eye: Self-Supervised Depth Reconstruction From Brain Activity
This paper presents the first method to reconstruct dense 3D depth maps directly from fMRI brain activity using self-supervised learning. By leveraging pre-trained monocular depth estimation on unpaired natural images and introducing a depth-based perceptual similarity metric, the approach outperforms indirect depth recovery from reconstructed images, with direct depth prediction achieving a mean depth-rank of 97 in a 1000-way identification task—significantly better than indirect methods.
In the past few years, significant advancements were made in reconstruction of observed natural images from fMRI brain recordings using deep-learning tools. Here, for the first time, we show that dense 3D depth maps of observed 2D natural images can also be recovered directly from fMRI brain recordings. We use an off-the-shelf method to estimate the unknown depth maps of natural images. This is applied to both: (i) the small number of images presented to subjects in an fMRI scanner (images for which we have fMRI recordings - referred to as "paired" data), and (ii) a very large number of natural images with no fMRI recordings ("unpaired data"). The estimated depth maps are then used as an auxiliary reconstruction criterion to train for depth reconstruction directly from fMRI. We propose two main approaches: Depth-only recovery and joint image-depth RGBD recovery. Because the number of available "paired" training data (images with fMRI) is small, we enrich the training data via self-supervised cycle-consistent training on many "unpaired" data (natural images & depth maps without fMRI). This is achieved using our newly defined and trained Depth-based Perceptual Similarity metric as a reconstruction criterion. We show that predicting the depth map directly from fMRI outperforms its indirect sequential recovery from the reconstructed images. We further show that activations from early cortical visual areas dominate our depth reconstruction results, and propose means to characterize fMRI voxels by their degree of depth-information tuning. This work adds an important layer of decoded information, extending the current envelope of visual brain decoding capabilities.
Motivation & Objective
- To extend fMRI-based visual decoding beyond RGB image reconstruction to include dense 3D depth maps.
- To address the scarcity of paired fMRI-image data by leveraging large-scale unpaired natural images with estimated depth maps.
- To develop a self-supervised training framework that improves depth reconstruction from fMRI using perceptual consistency.
- To characterize fMRI voxels by their sensitivity to depth information versus color/RGB content.
Proposed method
- Estimate depth maps for both paired fMRI images and a large set of unpaired natural images using a pre-trained monocular depth model (MiDaS).
- Train a deep encoder-decoder network to map fMRI activity to depth maps (or RGBD) using cycle-consistent self-supervision on unpaired data.
- Introduce a novel Depth-based Perceptual Similarity metric that uses features from a pre-trained depth network to guide reconstruction fidelity.
- Use a perceptually grounded loss function to ensure reconstructed depth maps are perceptually meaningful and consistent with human depth perception.
- Train two main architectures: depth-only recovery and joint RGBD recovery, with shared encoder and decoder components.
- Propose a Voxel Depth Sensitivity Index (VDSI) to quantify how selectively individual fMRI voxels encode depth versus RGB information.
Experimental results
Research questions
- RQ1Can dense 3D depth maps be reconstructed directly from fMRI brain activity, without relying on intermediate image reconstruction?
- RQ2Does self-supervised training on unpaired natural images with estimated depth maps improve fMRI-based depth reconstruction?
- RQ3How does direct depth prediction from fMRI compare to indirect depth recovery via reconstructed RGB images in terms of accuracy?
- RQ4Which brain regions and voxels are most sensitive to depth information in fMRI activity?
- RQ5Can individual fMRI voxels be systematically characterized by their degree of depth-information tuning?
Key findings
- The direct RGBD approach achieved a mean depth-rank of 97 in a 1000-way identification task, significantly outperforming indirect methods (38 and 56 rank levels respectively).
- Depth reconstruction performance is predominantly driven by early visual cortex (LVC, V1–V3) voxels, while higher visual cortex (HVC) voxels contribute less to depth decoding.
- The proposed Depth-based Perceptual Similarity metric effectively guides reconstruction, producing perceptually plausible depth maps.
- The Voxel Depth Sensitivity Index (VDSI) shows strong consistency across datasets, with most voxels exhibiting low sensitivity to depth compared to color.
- The direct depth prediction method achieves high-quality depth maps with a mean RGB-rank of 16 in the same 1000-way task, indicating strong joint image and depth reconstruction performance.
- The study demonstrates that depth is a viable and informative modality for fMRI-based brain decoding, extending the scope of visual information recoverable from brain activity.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.