[Paper Review] Molecular-driven Foundation Model for Oncologic Pathology
Threads is a slide-level foundation model that learns universal whole-slide embeddings by jointly encoding histology and molecular profiles, achieving state-of-the-art performance across 54 oncology tasks and enabling effective few-shot and transfer learning.
Foundation models are reshaping computational pathology by enabling transfer learning, where models pre-trained on vast datasets can be adapted for downstream diagnostic, prognostic, and therapeutic response tasks. Despite these advances, foundation models are still limited in their ability to encode the entire gigapixel whole-slide images without additional training and often lack complementary multimodal data. Here, we introduce Threads, a slide-level foundation model capable of generating universal representations of whole-slide images of any size. Threads was pre-trained using a multimodal learning approach on a diverse cohort of 47,171 hematoxylin and eosin (H&E)-stained tissue sections, paired with corresponding genomic and transcriptomic profiles - the largest such paired dataset to be used for foundation model development to date. This unique training paradigm enables Threads to capture the tissue's underlying molecular composition, yielding powerful representations applicable to a wide array of downstream tasks. In extensive benchmarking across 54 oncology tasks, including clinical subtyping, grading, mutation prediction, immunohistochemistry status determination, treatment response prediction, and survival prediction, Threads outperformed all baselines while demonstrating remarkable generalizability and label efficiency. It is particularly well suited for predicting rare events, further emphasizing its clinical utility. We intend to make the model publicly available for the broader community.
Motivation & Objective
- Address data scarcity and gigapixel WSI size by creating universal slide embeddings that generalize across cancer types and tasks.
- Leverage multimodal pretraining by aligning WSIs with accompanying molecular (transcriptomic and genomic) data to capture biologically relevant tissue information.
- Evaluate Threads across a broad benchmark (54 tasks, 23 cohorts) to demonstrate generalizability, transferability, and label efficiency.
- Provide analysis of data efficiency, transfer to external cohorts, and few-shot performance to support clinical utility.
Proposed method
- Two-part architecture: an ROI encoder (CONCHv1.5 ViT-L) for patches and a slide encoder with attention that aggregates tile embeddings into a slide representation.
- Multimodal pretraining via cross-modal contrastive learning to align WSI embeddings with corresponding molecular profiles (transcriptomics and targeted genomics).
- Training on MBTG-47k, a large histomolecular dataset (>47k samples) from MGH, BWH, TCGA, and GTEx.
- Downstream evaluation on 54 tasks across four families (clinical subtyping/grading, mutation prediction, IHC status, treatment/prognosis) using linear probing and various metrics.
- Optional fine-tuning: Threads initialized from pretraining can be fine-tuned for downstream tasks with notable gains.
- Introduction of molecular prompting to enable zero-shot-like classification using molecular prototypes.
Experimental results
Research questions
- RQ1Can a slide-level encoder learned with multimodal molecular guidance produce universal WSI embeddings applicable across diverse oncology tasks?
- RQ2How does Threads compare to existing whole-slide encoders (Prism, GigaPath, Chief) across a wide benchmark in terms of accuracy and robustness?
- RQ3What is the transferability of Threads embeddings to external cohorts and its performance in data-scarce, rare-event scenarios?
- RQ4Does Threads offer data- and label-efficient performance, including few-shot learning and effective fine-tuning, for clinically relevant tasks?
Key findings
- Threads achieves state-of-the-art performance on 54 tasks, outperforming Prism, GigaPath, and Chief with absolute gains of 6.3%, 9.9%, and 6.7% in linear probing respectively (P<0.001).
- Threads yields strong task-level gains, e.g., 2.1% improvement in clinical subtyping/grading, 6.1% in mutation prediction, 4.6% in IHC status, and 8.9% in prognostication (relative to best baselines).
- Threads demonstrates robust transferability to external cohorts, outperforming baselines in 8/9 transfer tasks and maintaining high AUCs in breast and lung subtyping (e.g., 98.4% and 96.5% AUC).
- In data-scarce settings, Threads shows superior few-shot performance and treatment-response/survival prediction especially for small-to-medium cohorts (absolute gains up to ~5–9% over baselines).
- Fine-tuning Threads gives substantial gains over training from scratch (average 2.2% across 54 tasks; up to 5.5% in mutation prediction).
- Threads enables molecular prompting, achieving competitive zero-shot-like performance using molecular prototypes for eight tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.