[論文レビュー] A Multimodal Knowledge-enhanced Whole-slide Pathology Foundation Model
本論文は mSTAR を導入する。WSIs、病理報告、RNA-Seq からの多模態知識を病理ファウンデーションモデルに注入する二段階の全スライド事前学習パラダイムであり、スライドレベルおよび多模態の下流タスクで優れた性能を発揮する。
Remarkable strides in computational pathology have been made in the task-agnostic foundation model that advances the performance of a wide array of downstream clinical tasks. Despite the promising performance, there are still several challenges. First, prior works have resorted to either vision-only or image-caption data, disregarding pathology reports with more clinically authentic information from pathologists and gene expression profiles which respectively offer distinct knowledge for versatile clinical applications. Second, the current progress in pathology FMs predominantly concentrates on the patch level, where the restricted context of patch-level pretraining fails to capture whole-slide patterns. Even recent slide-level FMs still struggle to provide whole-slide context for patch representation. In this study, for the first time, we develop a pathology foundation model incorporating three levels of modalities: pathology slides, pathology reports, and gene expression data, which resulted in 26,169 slide-level modality pairs from 10,275 patients across 32 cancer types, amounting to over 116 million pathological patch images. To leverage these data for CPath, we propose a novel whole-slide pretraining paradigm that injects the multimodal whole-slide context into the patch representation, called Multimodal Self-TAught PRetraining (mSTAR). The proposed paradigm revolutionizes the pretraining workflow for CPath, enabling the pathology FM to acquire the whole-slide context. To the best of our knowledge, this is the first attempt to incorporate three modalities at the whole-slide context for enhancing pathology FMs. To systematically evaluate the capabilities of mSTAR, we built the largest spectrum of oncological benchmark, spanning 7 categories of oncological applications in 15 types of 97 practical oncological tasks.
研究の動機と目的
- 多様なモーダルデータ(WSIs、病理報告、RNA-Seq)を病理ファウンデーションモデルに活用することを、ビジョンのみまたはキャプションベースのデータに限定しない動機付け。
- パッチレベルの事前学習の限界を、全スライドの文脈で知識を注入することで克服。
- スライドレベルの対照学習と自己教師付き学習を可能にする、TCGAベースの大規模な多模態データセットをキュレーション。
- スライドレベルの知識をパッチ抽出器へ転移させる二段階の事前学習パラダイム(スライドレベル対照学習; パッチレベル自己教師付き学習)を開発。
- 多くのスライドレベル診断、分子、予後、マルチモーダル融合タスクで改善を実証。
提案手法
- Stage 1: Slide-level contrastive learning to inject multimodal knowledge into a slide aggregator using WSIs, pathology reports, and RNA-Seq data.
- Stage 2: Self-taught training where the pretrained slide aggregator acts as a teacher to guide the patch extractor to reproduce slide-level embeddings.
- Use patch features from a pretrained extractor (UNI) fed into a slide aggregator for slide-level representation and inter-modality alignment.
- Incorporate inter-cancer contrastive learning to mitigate heterogeneity across cancer types.
- Evaluate with seven application types across 43 subtasks including unimodal and multimodal tasks.
- Assess the benefits of a pretrained aggregator (TransMIL+) with various patch extractors (e.g., mSTAR, UNI, CONCH, etc.).
実験結果
リサーチクエスチョン
- RQ1Can whole-slide multimodal knowledge improve pathology foundation models beyond patch-level, unimodal pretraining?
- RQ2Does slide-level contrastive learning with WSIs, pathology reports, and RNA-Seq enable better downstream performance across diagnostic, molecular, and prognostic tasks?
- RQ3Does self-taught training effectively transfer slide-level multimodal knowledge to patch extractors?
- RQ4How does incorporating a pretrained slide aggregator affect multimodal fusion and few-shot/zero-shot slide classification?
- RQ5What is the impact of multimodal pretraining on survival prognostics compared with unimodal baselines?
主な発見
- mSTAR achieves consistent performance gains across 12 slide-classification subtasks (diagnostic and molecular) versus patch-level SOTA models.
- mSTAR+ (with a pretrained aggregator) improves several diagnostic tasks, e.g., CAMELYON, with significant differences (P<0.001).
- In molecular prediction, mSTAR shows notable gains, including BRCA-Molecular and CRC-Molecular, with up to +4.60% improvements in certain tasks when using the aggregator.
- For survival analysis across 9 TCGA datasets, mSTAR (TransMIL) outperforms baselines with an average C-Index gain of +1.98% relative to UNI, and mSTAR+ yields further gains in several cancers.
- Multimodal fusion using mSTAR as patch features outperforms SOTA fusion models (MCAT, Porpoise, MOTCat, CMTA) by substantial margins across datasets.
- Few-shot slide classification with MI-FewShot shows mSTAR achieving the best overall rank across 6 subtasks, with notable gains in several tasks.
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。