Skip to main content
QUICK REVIEW

[Paper Review] Atlas: A Novel Pathology Foundation Model by Mayo Clinic, Charité, and Aignostics

Maximilian Alber, Stephan Tietz|arXiv (Cornell University)|Jan 9, 2025
Tuberculosis Research and Epidemiology3 citations
TL;DR

Atlas is a pathology foundation model trained on 1.2 million WSIs from Mayo Clinic and Charité, achieving state-of-the-art average performance across 21 public pathology benchmarks without being the largest model or dataset.

ABSTRACT

Recent advances in digital pathology have demonstrated the effectiveness of foundation models across diverse applications. In this report, we present Atlas, a novel vision foundation model based on the RudolfV approach. Our model was trained on a dataset comprising 1.2 million histopathology whole slide images, collected from two medical institutions: Mayo Clinic and Charité - Universtätsmedizin Berlin. Comprehensive evaluations show that Atlas achieves state-of-the-art performance across twenty-one public benchmark datasets, even though it is neither the largest model by parameter count nor by training dataset size.

Motivation & Objective

  • Motivate robust, generalizable representations for histopathology via large-scale self-supervised learning.
  • Leverage multi-stain, multi-magnification WSIs to cover diverse tissue types and scanner variations.
  • Evaluate Atlas across a broad suite of downstream pathology tasks to assess generalization.
  • Compare Atlas to other leading pathology foundation models to position its strengths and limitations.

Proposed method

  • Train a ViT-H/14 pathology foundation model (632M parameters) using an adapted RudolfV self-supervised approach based on DINOv2 framework.
  • Use a dataset of 1.2M de-identified WSIs from Mayo Clinic and Charité, with tiles generated at multiple resolutions (0.25, 0.5, 1.0, 2.0 µm/pixel).
  • Sample data to ~520M tiles for training; perform training on Nvidia H100 GPUs within Mayo Clinic Platform.
  • Evaluate embeddings via linear probing and ABMIL-style slide-level methods across 21 public benchmarks using both CLS and CLS+Mean token representations.
  • Assess performance with balanced accuracy for patch-level tasks and ABMIL-based slide-level tasks; report mean and standard errors over seeds.

Experimental results

Research questions

  • RQ1How does Atlas perform on a wide range of morphology- and molecular-related pathology tasks compared to existing foundation models?
  • RQ2Does multi-stain and multi-magnification training confer robustness and generalization advantages across diverse datasets and scanners?
  • RQ3What is the impact of the chosen token representation (CLS vs CLS+Mean) on downstream performance?
  • RQ4Can Atlas achieve state-of-the-art results without being the largest model or trained on the largest dataset?

Key findings

  • Atlas achieves an average performance of 61.9% across 21 benchmarks, outperforming Virchow2 and H-Optimus-0 by 1.1 percentage points on average.
  • Atlas shows the best performance on 11 of 21 benchmarks across molecular- and morphology-related tasks and is second-best on many others.
  • Across molecular-related tasks, Atlas ranks first on several HEST tasks and compounds overall performance with top-2 placements in many benchmarks.
  • In morphology-related benchmarks, Atlas delivers top performance on multiple datasets such as MSI CRC, MSI STAD, TCGA Uniform, BACH, CRC-100k, MHIST, PCAM, CAMELYON16, and PANDA.
  • Atlas's performance is close to or surpasses state-of-the-art models even though it is not the largest by parameters or data size, suggesting strong generalization from diverse training data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.