Skip to main content
QUICK REVIEW

[论文解读] Brain Diffusion for Visual Exploration: Cortical Discovery using Large Scale Generative Models

Andrew F. Luo, Margaret M. Henderson|arXiv (Cornell University)|Jun 5, 2023
Cell Image Analysis Techniques被引用 6
一句话总结

BrainDiVE 提出了一种数据驱动的生成框架,利用大规模扩散模型并以 fMRI 脑活动为指导,生成可预测地激活特定视觉皮层区域的图像。通过用脑激活最大化替代文本提示,该方法揭示了精细的功能特化,识别出类别选择性 ROI 内的新亚区,并通过行为研究验证了结果,推动了对人类视觉皮层组织的无假设探索。

ABSTRACT

A long standing goal in neuroscience has been to elucidate the functional organization of the brain. Within higher visual cortex, functional accounts have remained relatively coarse, focusing on regions of interest (ROIs) and taking the form of selectivity for broad categories such as faces, places, bodies, food, or words. Because the identification of such ROIs has typically relied on manually assembled stimulus sets consisting of isolated objects in non-ecological contexts, exploring functional organization without robust a priori hypotheses has been challenging. To overcome these limitations, we introduce a data-driven approach in which we synthesize images predicted to activate a given brain region using paired natural images and fMRI recordings, bypassing the need for category-specific stimuli. Our approach -- Brain Diffusion for Visual Exploration ("BrainDiVE") -- builds on recent generative methods by combining large-scale diffusion models with brain-guided image synthesis. Validating our method, we demonstrate the ability to synthesize preferred images with appropriate semantic specificity for well-characterized category-selective ROIs. We then show that BrainDiVE can characterize differences between ROIs selective for the same high-level category. Finally we identify novel functional subdivisions within these ROIs, validated with behavioral data. These results advance our understanding of the fine-grained functional organization of human visual cortex, and provide well-specified constraints for further examination of cortical organization using hypothesis-driven methods.

研究动机与目标

  • 为克服传统手工设计刺激在绘制高级视觉皮层功能组织时的局限性。
  • 利用基于真实 fMRI 数据的生成模型,实现数据驱动、无假设的皮层选择性探索。
  • 识别已知类别选择性 ROI 内的精细功能特化及新亚区。
  • 通过行为实验验证合成图像的语义,确保其生态和感知相关性。
  • 提供一种新的、客观的方法,利用大规模生成式 AI 发现人类视觉皮层的功能组织。

提出的方法

  • 该方法使用预训练的扩散模型,其条件不是基于文本提示,而是基于特定脑区的 fMRI 体素活动模式。
  • 将脑引导的图像生成建模为梯度优化问题,通过 CLIP 嵌入对齐使图像更新最大化目标脑区的激活。
  • 模型基于 CLIP 图像嵌入对 fMRI 响应进行线性拟合所得的每个体素权重向量的平均值计算梯度,从而有效利用脑信号作为条件信号。
  • 优化过程迭代地改进生成图像,使其 CLIP 嵌入方向与目标脑区体素权重的联合方向对齐。
  • 通过优化扩散过程,使生成图像同时满足自然图像先验和脑激活约束。
  • 该方法利用大规模自然场景图像的 fMRI 数据集,实现真实且生态有效的刺激生成。
Figure 1: Images generated using BrainDiVE . Images are generated using a diffusion model with maximization of voxels identified from functional localizer experiments as conditioning. We find that brain signals recorded via fMRI can guide the synthesis of images with high semantic specificity, stren
Figure 1: Images generated using BrainDiVE . Images are generated using a diffusion model with maximization of voxels identified from functional localizer experiments as conditioning. We find that brain signals recorded via fMRI can guide the synthesis of images with high semantic specificity, stren

实验结果

研究问题

  • RQ1fMRI 引导的扩散模型能否生成具有高语义特异性的图像,使其被预测为选择性激活已知的类别选择性视觉皮层区域?
  • RQ2该方法能否揭示对同一高层类别(如人脸或场景)选择的脑区之间的功能差异?
  • RQ3该方法能否识别出在传统刺激集下无法检测到的既存类别选择性 ROI 内的新功能亚区?
  • RQ4合成图像是否能在人类观察者中引发一致的感知反应,从而确认其语义身份和生态相关性?
  • RQ5脑引导的图像生成在多大程度上能提供关于人类视觉皮层精细功能架构的新、可检验的假设?

主要发现

  • BrainDiVE 生成了具有高度语义特异性的图像,成功激活了已知的类别选择性区域,如梭状回面孔区(fusiform face area)和海马旁位置区(parahippocampal place area)。
  • 该方法揭示了同一高层类别内部的功能特化差异,例如在对象构成和空间上下文敏感性方面,面孔选择性与场景选择性区域之间存在差异。
  • 在现有 ROI 内识别出新的功能亚区,例如腹侧颞叶中的某些亚区对特定视觉特征表现出差异性敏感性,超越了宽泛的类别标签。
  • 行为研究证实,人类受试者能正确分类合成图像的语义内容,验证了其感知和生态相关性。
  • 该方法在多个受试者和脑区中表现出鲁棒性,每张图像的图像生成耗时 20–30 秒(在 V100 GPU 上)。
  • 该方法消耗了约 1,500 GPU 小时,表明其在大规模皮层探索中的可扩展性。
Figure 2: Architecture of brain guided diffusion (BrainDiVE) . Top: Our framework consists of two core components: (1) A diffusion model trained to synthesize natural images by iterative denoising; we utilize pretrained LDMs. (2) An encoder trained to map from images to cortical activity. Our framew
Figure 2: Architecture of brain guided diffusion (BrainDiVE) . Top: Our framework consists of two core components: (1) A diffusion model trained to synthesize natural images by iterative denoising; we utilize pretrained LDMs. (2) An encoder trained to map from images to cortical activity. Our framew

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。