Skip to main content
QUICK REVIEW

[论文解读] Deep Learning Predicts Biomarker Status and Discovers Related Histomorphology Characteristics for Low-Grade Glioma

Zijie Fang, Yihan Liu|arXiv (Cornell University)|Oct 11, 2023
Digital Imaging for Blood Diseases被引用 4
一句话总结

本文提出 Multi-Beholder,一种深度学习流程,仅使用苏木精-伊红染色(H&E)全切片图像即可预测五种低级别胶质瘤(LGG)生物标志物,结合多实例学习与单类分类以实现准确的实例级伪标签生成。在 TCGA-LGG 数据集上 AUC 最高达 0.973,在外部队列中达 0.820,实现了对生物标志物状态相关组织形态学特征的可解释性发现。

ABSTRACT

Biomarker detection is an indispensable part in the diagnosis and treatment of low-grade glioma (LGG). However, current LGG biomarker detection methods rely on expensive and complex molecular genetic testing, for which professionals are required to analyze the results, and intra-rater variability is often reported. To overcome these challenges, we propose an interpretable deep learning pipeline, a Multi-Biomarker Histomorphology Discoverer (Multi-Beholder) model based on the multiple instance learning (MIL) framework, to predict the status of five biomarkers in LGG using only hematoxylin and eosin-stained whole slide images and slide-level biomarker status labels. Specifically, by incorporating the one-class classification into the MIL framework, accurate instance pseudo-labeling is realized for instance-level supervision, which greatly complements the slide-level labels and improves the biomarker prediction performance. Multi-Beholder demonstrates superior prediction performance and generalizability for five LGG biomarkers (AUROC=0.6469-0.9735) in two cohorts (n=607) with diverse races and scanning protocols. Moreover, the excellent interpretability of Multi-Beholder allows for discovering the quantitative and qualitative correlations between biomarker status and histomorphology characteristics. Our pipeline not only provides a novel approach for biomarker prediction, enhancing the applicability of molecular treatments for LGG patients but also facilitates the discovery of new mechanisms in molecular functionality and LGG progression.

研究动机与目标

  • 开发一种成本低廉、可解释的方法,以实现无需依赖昂贵分子检测的 LGG 生物标志物预测。
  • 克服当前生物标志物检测的局限性,包括高成本、技术复杂性以及阅片者间差异性。
  • 通过利用常规 H&E 全切片图像,推动分子诊断在临床中的广泛应用。
  • 发现与 LGG 中特定生物标志物状态相关的定量与定性组织形态学特征。
  • 利用深度学习提高低级别胶质瘤中生物标志物预测的可及性与准确性。

提出的方法

  • 该框架将多实例学习(MIL)与单类分类相结合,利用切片级标注生成准确的实例级伪标签。
  • 对每个全切片图像应用单类分类,以识别最可能反映生物标志物状态的代表性组织切片。
  • 模型仅使用切片级标签学习区分正负样本的切片级特征,从而减少对精确切片级标注的依赖。
  • 双分支注意力机制通过聚焦于多个组织切片中形态学相关的区域,增强特征学习。
  • 该流程在全切片图像上端到端训练,推理在切片级别进行,以预测生物标志物状态。
  • 通过注意力图与显著性可视化实现可解释性,将预测的生物标志物状态与特定组织学模式相关联。

实验结果

研究问题

  • RQ1仅使用 H&E 染色的全切片图像,深度学习模型能否准确预测五种关键 LGG 生物标志物?
  • RQ2在多实例学习框架中,单类分类如何提升生物标志物预测的伪标签生成准确性?
  • RQ3哪些组织形态学特征在定量与定性上与 LGG 中特定生物标志物状态相关?
  • RQ4该模型在不同患者人群与扫描协议下的泛化能力如何?
  • RQ5模型的可解释性能否揭示与分子功能与肿瘤进展相关的生物学上有意义的模式?

主要发现

  • 在内部 TCGA-LGG 队列中,模型 AUC 达 0.973,表明其具有极高的预测性能。
  • 在外部分离的湘雅队列(涵盖不同种族与扫描协议)中,模型 AUC 达 0.820,表明其具备强大的泛化能力。
  • 单类分类的整合显著提升了实例级伪标签生成的准确性,从而增强了切片级预测性能。
  • 可解释性分析揭示了与每种生物标志物相关联的显著组织形态学模式,包括核密度、染色质纹理与细胞排列方式。
  • 该流程成功发现了与分子功能相关的新型形态学关联,提示可能揭示 LGG 进展的生物学机制。
  • 模型在不同扫描协议与患者人口统计学特征下表现稳健,支持其临床转化潜力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。