Skip to main content
QUICK REVIEW

[论文解读] Sharing deep generative representation for perceived image reconstruction from human brain activity

Changde Du, Changying Du|arXiv (Cornell University)|Apr 25, 2017
Cell Image Analysis Techniques被引用 15
一句话总结

该论文提出了一种深度生成多视图模型(DGMM),通过共享的非线性潜在表示联合建模刺激和神经响应,从fMRI脑活动重建感知的视觉图像。通过利用深度神经网络进行图像建模,采用全秩协方差表示fMRI体素相关性,并结合变分贝叶斯推断与后验正则化,DGMM在三个数据集上均实现了最先进的重建性能,显著优于先前方法,在SSIM和PCC指标上表现更优。

ABSTRACT

Decoding human brain activities via functional magnetic resonance imaging (fMRI) has gained increasing attention in recent years. While encouraging results have been reported in brain states classification tasks, reconstructing the details of human visual experience still remains difficult. Two main challenges that hinder the development of effective models are the perplexing fMRI measurement noise and the high dimensionality of limited data instances. Existing methods generally suffer from one or both of these issues and yield dissatisfactory results. In this paper, we tackle this problem by casting the reconstruction of visual stimulus as the Bayesian inference of missing view in a multiview latent variable model. Sharing a common latent representation, our joint generative model of external stimulus and brain response is not only "deep" in extracting nonlinear features from visual images, but also powerful in capturing correlations among voxel activities of fMRI recordings. The nonlinearity and deep structure endow our model with strong representation ability, while the correlations of voxel activities are critical for suppressing noise and improving prediction. We devise an efficient variational Bayesian method to infer the latent variables and the model parameters. To further improve the reconstruction accuracy, the latent representations of testing instances are enforced to be close to that of their neighbours from the training set via posterior regularization. Experiments on three fMRI recording datasets demonstrate that our approach can more accurately reconstruct visual stimuli.

研究动机与目标

  • 为解决从噪声大、高维的fMRI数据中重建详细视觉刺激的挑战。
  • 克服现有fMRI到图像重建方法中线性模型的局限性以及噪声抑制能力差的问题。
  • 开发一种深层生成模型,以捕捉图像中的非线性特征以及fMRI体素之间的相关性。
  • 通过在测试阶段潜在表示上施加后验正则化,强制结构一致性,从而提高重建精度。
  • 提供一种概率性、贝叶斯框架,通过模型平均减少小样本fMRI数据集上的过拟合。

提出的方法

  • 将图像重建问题建模为多视图潜在变量模型中缺失视图的贝叶斯推断,该模型具有共享潜在空间。
  • 使用深度神经网络(如CNN或MLP)作为视觉图像的非线性观测模型,以增强特征表示能力。
  • 采用全秩高斯分布建模fMRI体素活动,通过降低为低秩协方差以控制计算成本,同时保留体素间相关性。
  • 采用高效的平均场变分贝叶斯推断方法估计潜在变量和模型参数。
  • 应用后验正则化,约束测试阶段的潜在表示与训练集邻居的潜在表示保持接近,从而提升泛化能力。
  • 通过贝叶斯推断实现模型平均,以缓解小样本fMRI数据上的过拟合问题。

实验结果

研究问题

  • RQ1具有共享潜在表示的深度生成模型是否能在从fMRI重建视觉图像方面优于线性模型或非生成模型?
  • RQ2fMRI体素活动的全秩协方差结构在多大程度上能提升噪声抑制能力和重建精度?
  • RQ3后验正则化在将测试阶段潜在表示与训练集邻居对齐方面有多有效,从而提升泛化能力?
  • RQ4在图像建模中引入非线性深度特征是否能带来优于线性或浅层模型的重建效果?
  • RQ5在小样本fMRI数据集上,完全生成的贝叶斯框架是否能实现优于判别式或两阶段方法的性能?

主要发现

  • 在数据集1上,DGMM的PCC达到0.611 ± 0.183,显著优于BCCA(0.438 ± 0.215)和DCCAE变体。
  • 在数据集2上,DGMM实现了最高的SSIM(0.645 ± 0.054),超过De-CNN(0.613 ± 0.043)及其他基线方法。
  • 在数据集3上,DGMM在重建图像上的分类准确率达到0.778 ± 0.083,远高于BCCA(0.600 ± 0.098)和DCCAE-S(0.478 ± 0.155)。
  • DGMM在所有三个数据集上的SSIM均表现出一致优势,尤其在数据集2上较次优方法提升高达0.15。
  • 该模型在感知相似性(SSIM)和相关性(PCC)两个指标上均达到最先进水平,表明重建保真度显著提升。
  • 后验正则化有助于提升泛化能力,表现为在测试样本上持续获得性能增益。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。