Skip to main content
QUICK REVIEW

[论文解读] Pathomic Fusion: An Integrated Framework for Fusing Histopathology and Genomic Features for Cancer Diagnosis and Prognosis

Richard J. Chen, Ming Y. Lu|arXiv (Cornell University)|Dec 18, 2019
AI in cancer detection被引用 11
一句话总结

Pathomic Fusion 提出了一种新颖的、可解释的多模态深度学习框架,通过融合组织病理学图像与多组学数据(突变、CNV、RNA-Seq)以改善癌症预后和诊断。该模型采用基于门控的注意力机制控制模态表达能力,并通过克罗内克积实现特征交互,优于单模态网络,在 TCGA 的胶质瘤和透明细胞肾细胞癌数据集中,通过图像和细胞图上的显著性图提升了可解释性,实现了更优的生存预测性能。

ABSTRACT

Cancer diagnosis, prognosis, and therapeutic response predictions are based on morphological information from histology slides and molecular profiles from genomic data. However, most deep learning-based objective outcome prediction and grading paradigms are based on histology or genomics alone and do not make use of the complementary information in an intuitive manner. In this work, we propose Pathomic Fusion, an interpretable strategy for end-to-end multimodal fusion of histology image and genomic (mutations, CNV, RNA-Seq) features for survival outcome prediction. Our approach models pairwise feature interactions across modalities by taking the Kronecker product of unimodal feature representations and controls the expressiveness of each representation via a gating-based attention mechanism. Following supervised learning, we are able to interpret and saliently localize features across each modality, and understand how feature importance shifts when conditioning on multimodal input. We validate our approach using glioma and clear cell renal cell carcinoma datasets from the Cancer Genome Atlas (TCGA), which contains paired whole-slide image, genotype, and transcriptome data with ground truth survival and histologic grade labels. In a 15-fold cross-validation, our results demonstrate that the proposed multimodal fusion paradigm improves prognostic determinations from ground truth grading and molecular subtyping, as well as unimodal deep networks trained on histology and genomic data alone. The proposed method establishes insight and theory on how to train deep networks on multimodal biomedical data in an intuitive manner, which will be useful for other problems in medicine that seek to combine heterogeneous data streams for understanding diseases and predicting response and resistance to treatment.

研究动机与目标

  • 开发一种可解释的、端到端的框架,用于融合组织病理学图像与多组学生物特征,以实现癌症预后和诊断。
  • 解决单模态深度学习模型无法利用组织学与基因组学之间互补信息的局限性。
  • 通过基于注意力的可解释性,实现在图像和基因组模态上关键特征的显著定位。
  • 理解在多模态输入条件下特征重要性如何变化,揭示形态学与分子特征之间的协同作用。
  • 提供一种可扩展且通用的范式,用于在肿瘤学中整合异构的生物医学数据流。

提出的方法

  • 使用卷积神经网络(CNNs)和图卷积网络(GCNs)分别从全切片图像和细胞图中提取形态学特征。
  • 在融合前,对单模态特征提取器(图像使用 CNNs,基因组使用 SNNs)采用独立的监督训练流程。
  • 应用基于门控的注意力机制控制各模态特征表示的表达能力,实现模态贡献的动态调制。
  • 通过门控特征表示的克罗内克积执行多模态融合,以建模模态间的成对交互。
  • 在 TCGA 的胶质瘤和 CCRCC 数据集上,通过交叉验证进行端到端训练,以预测生存结局。
  • 在图像和细胞图输入上使用基于梯度的显著性图(如 Grad-CAM)对预后预测中的显著区域和基因进行定位与解释。

实验结果

研究问题

  • RQ1能否通过融合组织病理学与多组学数据的多模态深度学习框架,在生存结局预测上超越单模态模型?
  • RQ2在多模态输入条件下,特征重要性评分如何变化?这揭示了形态学与分子谱型之间协同作用的哪些见解?
  • RQ3该模型能否在组织病理学图像中定位生物上相关的区域,以及在基因组数据中识别出对生存和组织学分级具有预测性的特定基因?
  • RQ4形态学与分子特征的整合在多大程度上能细化现有的分子亚型并改善患者分层?
  • RQ5注意力机制与克罗内克融合策略能否提供可解释的、具有临床意义的生物标志物,这些标志物在单模态分析中无法被检测到?

主要发现

  • Pathomic Fusion 在 TCGA 胶质瘤和 CCRCC 数据集的 15 折交叉验证中,显著优于仅基于组织学或基因组学训练的单模态模型,提升了生存预测性能。
  • 该模型成功定位了胶质瘤中的关键组织病理学特征,如微血管增生、肿瘤细胞密度和异型性,以及少突胶质瘤中的“蛋样细胞”,与已知的组织学模式一致。
  • 在胶质母细胞瘤中,模型识别出已知的预后相关基因如 PTEN、MYC、CDKN2A、EGFR 和 FGFR2 为高度重要,而 ANO9 和 RB1 的重要性在不同亚型中发生变化,与已知的肿瘤抑制基因和癌基因功能一致。
  • 对于 CCRCC,模型突出显示了 CYP3A7 表达降低和 PITX2、DDX43 与 XIST 表达增加为风险相关特征,而 HAGHL、MMP1 和 ARRP21 的重要性在结合形态学特征时上升。
  • 形态学与基因组特征的整合导致基因归因发生变化,例如在 IDH 野生型胶质母细胞瘤中,EGFR 扩增的重要性降低,支持其在该亚型中治疗相关性有限。
  • 该模型表明,细胞图可揭示在原始组织病理学图像中难以识别的显著生物相关细胞(如异型细胞),从而增强可解释性与临床相关性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。