Skip to main content
QUICK REVIEW

[论文解读] A Multimodal Knowledge-enhanced Whole-slide Pathology Foundation Model

Yingxue Xu, Yihui Wang|arXiv (Cornell University)|Jul 22, 2024
Biomedical Text Mining and Ontologies被引用 9
一句话总结

本论文提出 mSTAR,是一个两阶段的全切片预训练范式,从 WSIs、病理报告和 RNA-Seq 膜注入多模态知识到病理 foundation 模型,使得切片级及多模态下游任务具备更优性能。

ABSTRACT

Remarkable strides in computational pathology have been made in the task-agnostic foundation model that advances the performance of a wide array of downstream clinical tasks. Despite the promising performance, there are still several challenges. First, prior works have resorted to either vision-only or image-caption data, disregarding pathology reports with more clinically authentic information from pathologists and gene expression profiles which respectively offer distinct knowledge for versatile clinical applications. Second, the current progress in pathology FMs predominantly concentrates on the patch level, where the restricted context of patch-level pretraining fails to capture whole-slide patterns. Even recent slide-level FMs still struggle to provide whole-slide context for patch representation. In this study, for the first time, we develop a pathology foundation model incorporating three levels of modalities: pathology slides, pathology reports, and gene expression data, which resulted in 26,169 slide-level modality pairs from 10,275 patients across 32 cancer types, amounting to over 116 million pathological patch images. To leverage these data for CPath, we propose a novel whole-slide pretraining paradigm that injects the multimodal whole-slide context into the patch representation, called Multimodal Self-TAught PRetraining (mSTAR). The proposed paradigm revolutionizes the pretraining workflow for CPath, enabling the pathology FM to acquire the whole-slide context. To the best of our knowledge, this is the first attempt to incorporate three modalities at the whole-slide context for enhancing pathology FMs. To systematically evaluate the capabilities of mSTAR, we built the largest spectrum of oncological benchmark, spanning 7 categories of oncological applications in 15 types of 97 practical oncological tasks.

研究动机与目标

  • Motivate leveraging multimodal data (WSIs, pathology reports, RNA-Seq) for pathology foundation models rather than vision-only or caption-based data.
  • Overcome patch-level pretraining limitations by injecting knowledge at the whole-slide context.
  • Curate a large multimodal TCGA-based dataset to enable slide-level contrastive pretraining and self-taught training.
  • Develop a two-stage pretraining paradigm (slide-level contrastive learning; patch-level self-taught training) to transfer slide-level knowledge to patch extractors.
  • Demonstrate improvements across a wide set of slide-level diagnostic, molecular, prognostic, and multimodal fusion tasks.

提出的方法

  • Stage 1: Slide-level contrastive learning to inject multimodal knowledge into a slide aggregator using WSIs, pathology reports, and RNA-Seq data.
  • Stage 2: Self-taught training where the pretrained slide aggregator acts as a teacher to guide the patch extractor to reproduce slide-level embeddings.
  • Use patch features from a pretrained extractor (UNI) fed into a slide aggregator for slide-level representation and inter-modality alignment.
  • Incorporate inter-cancer contrastive learning to mitigate heterogeneity across cancer types.
  • Evaluate with seven application types across 43 subtasks including unimodal and multimodal tasks.
  • Assess the benefits of a pretrained aggregator (TransMIL+) with various patch extractors (e.g., mSTAR, UNI, CONCH, etc.).

实验结果

研究问题

  • RQ1Can whole-slide multimodal knowledge improve pathology foundation models beyond patch-level, unimodal pretraining?
  • RQ2Does slide-level contrastive learning with WSIs, pathology reports, and RNA-Seq enable better downstream performance across diagnostic, molecular, and prognostic tasks?
  • RQ3Does self-taught training effectively transfer slide-level multimodal knowledge to patch extractors?
  • RQ4How does incorporating a pretrained slide aggregator affect multimodal fusion and few-shot/zero-shot slide classification?
  • RQ5What is the impact of multimodal pretraining on survival prognostics compared with unimodal baselines?

主要发现

  • mSTAR achieves consistent performance gains across 12 slide-classification subtasks (diagnostic and molecular) versus patch-level SOTA models.
  • mSTAR+ (with a pretrained aggregator) improves several diagnostic tasks, e.g., CAMELYON, with significant differences (P<0.001).
  • In molecular prediction, mSTAR shows notable gains, including BRCA-Molecular and CRC-Molecular, with up to +4.60% improvements in certain tasks when using the aggregator.
  • For survival analysis across 9 TCGA datasets, mSTAR (TransMIL) outperforms baselines with an average C-Index gain of +1.98% relative to UNI, and mSTAR+ yields further gains in several cancers.
  • Multimodal fusion using mSTAR as patch features outperforms SOTA fusion models (MCAT, Porpoise, MOTCat, CMTA) by substantial margins across datasets.
  • Few-shot slide classification with MI-FewShot shows mSTAR achieving the best overall rank across 6 subtasks, with notable gains in several tasks.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。