[论文解读] Interpretable multimodal fusion networks reveal mechanisms of brain cognition
该论文提出gCAM-CCL,一种可解释的多模态融合模型,通过结合基于梯度的类别特定激活图与协同学习,实现脑疾病分类与生物机制解释的同步进行。该方法识别出与认知表现相关的关键脑区和遗传标志物,揭示高WRAT表现者表现出更强的神经递质传递信号,而低表现者则因基因变异导致发育缺陷。
Multimodal fusion benefits disease diagnosis by providing a more comprehensive perspective. Developing algorithms is challenging due to data heterogeneity and the complex within- and between-modality associations. Deep-network-based data-fusion models have been developed to capture the complex associations and the performance in diagnosis has been improved accordingly. Moving beyond diagnosis prediction, evaluation of disease mechanisms is critically important for biomedical research. Deep-network-based data-fusion models, however, are difficult to interpret, bringing about difficulties for studying biological mechanisms. In this work, we develop an interpretable multimodal fusion model, namely gCAM-CCL, which can perform automated diagnosis and result interpretation simultaneously. The gCAM-CCL model can generate interpretable activation maps, which quantify pixel-level contributions of the input features. This is achieved by combining intermediate feature maps using gradient-based weights. Moreover, the estimated activation maps are class-specific, and the captured cross-data associations are interest/label related, which further facilitates class-specific analysis and biological mechanism analysis. We validate the gCAM-CCL model on a brain imaging-genetic study, and show gCAM-CCL's performed well for both classification and mechanism analysis. Mechanism analysis suggests that during task-fMRI scans, several object recognition related regions of interests (ROIs) are first activated and then several downstream encoding ROIs get involved. Results also suggest that the higher cognition performing group may have stronger neurotransmission signaling while the lower cognition performing group may have problem in brain/neuron development, resulting from genetic variations.
研究动机与目标
- 解决基于深度神经网络的多模态融合模型在脑疾病诊断中可解释性不足的问题。
- 开发一种方法,通过生成类别特定的像素级激活图,实现准确诊断与生物机制分析的同步进行。
- 利用多模态影像与基因组数据,识别与认知表现相关的脑功能连接模式与基因变异。
- 通过可解释的深度学习揭示认知能力差异的神经生物学机制。
- 在真实世界的阅读能力(WRAT)分类影像-基因组研究中验证该模型。
提出的方法
- gCAM-CCL 将基于梯度的类别激活映射(Grad-CAM)与协同学习层结合,从多模态输入(如fMRI和基因数据)生成类别特定的像素级激活图。
- 通过特征图的梯度反向传播计算重要性权重,再进行全局平均,生成可解释的激活图。
- 引入协同学习层,确保捕捉到的跨数据关联与目标特征(如高 vs. 低WRAT表现)相关,从而增强可解释性。
- 通过具有注意力类似加权的深度神经网络架构,融合功能磁共振成像(fMRI)衍生的脑功能连接(FC)矩阵与单核苷酸多态性(SNP)数据。
- 利用类别特定的激活图识别主导脑区与ROI-ROI连接,实现疾病或特征相关神经环路的可视化。
- 使用ConsensusPathDB-human数据库对模型识别出的SNP进行基因富集分析,将遗传标志物与生物通路关联。
实验结果
研究问题
- RQ1在阅读任务中,哪些脑区与功能连接模式最能预测高阶与低阶认知表现?
- RQ2如何使深度多模态融合模型具备可解释性,以揭示超越诊断准确性的生物意义机制?
- RQ3通过可解释的融合模型,哪些基因变异与生物通路与认知表现差异相关?
- RQ4所识别的功能连接模式与遗传标志物是否与已知的认知相关神经生物学过程一致?
- RQ5该模型能否基于影像与基因数据,同时区分高阶与低阶认知表现群体,并提供对潜在机制的可解释洞察?
主要发现
- 高WRAT组与三个枕叶区域——楔叶、中枕回与下枕回——之间的功能连接占主导地位,这些区域在物体与视觉词识别中起关键作用。
- 下游区域(如楔前叶与海马旁回)在初始视觉处理后被激活,表明任务中存在顺序处理级联。
- 低WRAT组未表现出主导枢纽ROI,但任务负性区域(如颞顶与扣带回)活动增强,提示认知处理网络激活较弱。
- 基因富集分析显示,高WRAT组的SNP显著与神经递质传递通路相关,包括突触信号传导与神经递质水平调节。
- 相反,低WRAT组的SNP在神经发育通路中富集,如中脑发育与生长锥导向,提示可能存在脑成熟缺陷。
- 该模型在分类低WRAT个体时特异性高于敏感性,表明识别出的低组功能连接模式区分度较低,可能更具噪声。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。