Skip to main content
QUICK REVIEW

[Paper Review] Multimodal Whole Slide Foundation Model for Pathology

Tong Ding, Sophia J. Wagner|arXiv (Cornell University)|Nov 29, 2024
Radiomics and Machine Learning in Medical Imaging23 citations
TL;DR

TITAN is a multimodal whole-slide foundation model pretrained on 335,645 whole-slide images with visual self-supervised learning and vision-language alignment to pathology reports and 423,122 synthetic captions, enabling zero-shot, few-shot, and cross-modal tasks without fine-tuning.

ABSTRACT

The field of computational pathology has been transformed with recent advances in foundation models that encode histopathology region-of-interests (ROIs) into versatile and transferable feature representations via self-supervised learning (SSL). However, translating these advancements to address complex clinical challenges at the patient and slide level remains constrained by limited clinical data in disease-specific cohorts, especially for rare clinical conditions. We propose TITAN, a multimodal whole slide foundation model pretrained using 335,645 WSIs via visual self-supervised learning and vision-language alignment with corresponding pathology reports and 423,122 synthetic captions generated from a multimodal generative AI copilot for pathology. Without any finetuning or requiring clinical labels, TITAN can extract general-purpose slide representations and generate pathology reports that generalize to resource-limited clinical scenarios such as rare disease retrieval and cancer prognosis. We evaluate TITAN on diverse clinical tasks and find that TITAN outperforms both ROI and slide foundation models across machine learning settings such as linear probing, few-shot and zero-shot classification, rare cancer retrieval and cross-modal retrieval, and pathology report generation.

Motivation & Objective

  • Develop a general-purpose, multimodal foundation model for pathology that operates at the slide level rather than ROI level.
  • Leverage large-scale WSI data with visual SSL and vision-language alignment to pathology reports.
  • Incorporate synthetic captions to enhance multimodal understanding without clinical labels.
  • Demonstrate generalization in rare disease retrieval, cancer prognosis, and cross-modal tasks without finetuning.

Proposed method

  • Pretrain TITAN on 335,645 whole-slide images using visual self-supervised learning.
  • Align visual representations with corresponding pathology reports in a vision-language objective.
  • Incorporate 423,122 synthetic captions generated by a multimodal AI copilot for pathology.
  • Evaluate zero-shot, few-shot, and linear probing on diverse clinical tasks without finetuning.

Experimental results

Research questions

  • RQ1Can a multimodal whole-slide foundation model derived from large-scale WSIs and synthetic captions generalize to resource-limited clinical settings without fine-tuning?
  • RQ2How does TITAN perform on rare cancer retrieval, cross-modal retrieval, and pathology report generation compared with ROI- and slide-based foundation models?

Key findings

  • TITAN outperforms ROI- and slide-foundation models across linear probing, few-shot, and zero-shot classification.
  • TITAN excels in rare cancer retrieval and cross-modal retrieval tasks.
  • TITAN can generate pathology reports and generalize to resource-limited clinical scenarios without requiring clinical labels.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.