[Paper Review] ACROBAT -- a multi-stain breast cancer histological whole-slide-image data set from routine diagnostics for computational pathology
ACROBAT is a large-scale, publicly available whole-slide image (WSI) dataset comprising 4,212 matched H&E and IHC-stained WSIs from 1,153 primary breast cancer patients, collected from routine diagnostics. It enables computational pathology research in image registration, stain-guided learning, virtual staining, and artifact detection, with performance benchmarking via 37,000 manually annotated landmark pairs from 13 annotators.
The analysis of FFPE tissue sections stained with haematoxylin and eosin (H&E) or immunohistochemistry (IHC) is an essential part of the pathologic assessment of surgically resected breast cancer specimens. IHC staining has been broadly adopted into diagnostic guidelines and routine workflows to manually assess status and scoring of several established biomarkers, including ER, PGR, HER2 and KI67. However, this is a task that can also be facilitated by computational pathology image analysis methods. The research in computational pathology has recently made numerous substantial advances, often based on publicly available whole slide image (WSI) data sets. However, the field is still considerably limited by the sparsity of public data sets. In particular, there are no large, high quality publicly available data sets with WSIs of matching IHC and H&E-stained tissue sections. Here, we publish the currently largest publicly available data set of WSIs of tissue sections from surgical resection specimens from female primary breast cancer patients with matched WSIs of corresponding H&E and IHC-stained tissue, consisting of 4,212 WSIs from 1,153 patients. The primary purpose of the data set was to facilitate the ACROBAT WSI registration challenge, aiming at accurately aligning H&E and IHC images. For research in the area of image registration, automatic quantitative feedback on registration algorithm performance remains available through the ACROBAT challenge website, based on more than 37,000 manually annotated landmark pairs from 13 annotators. Beyond registration, this data set has the potential to enable many different avenues of computational pathology research, including stain-guided learning, virtual staining, unsupervised pre-training, artefact detection and stain-independent models.
Motivation & Objective
- To address the scarcity of large, high-quality public WSI datasets with matched H&E and IHC stains from routine clinical diagnostics.
- To support the development and benchmarking of image registration algorithms for aligning H&E and IHC whole-slide images.
- To enable diverse computational pathology research, including stain-guided learning, virtual staining, and stain-independent model training.
- To provide a standardized benchmark for registration performance using 37,000 manually annotated landmark pairs.
Proposed method
- Collection of whole-slide images from surgical resection specimens of female primary breast cancer patients across multiple clinical sites.
- Acquisition of matched H&E and IHC-stained WSIs for each patient, preserving diagnostic quality and clinical relevance.
- Manual annotation of 37,000 landmark pairs across 13 annotators to enable quantitative evaluation of image registration algorithms.
- Publication of the dataset via arXiv and the ACROBAT challenge website to ensure accessibility and reproducibility.
- Design of the dataset to support multiple computational pathology tasks beyond registration, including unsupervised pre-training and artifact detection.
- Standardized data curation and quality control to ensure consistency and reliability for research applications.
Experimental results
Research questions
- RQ1Can deep learning models achieve accurate and robust registration between H&E and IHC whole-slide images using this dataset?
- RQ2To what extent can stain-guided learning improve model generalization across different staining protocols?
- RQ3How effective are virtual staining methods when trained on this multi-stain, clinically acquired WSI dataset?
- RQ4Can unsupervised pre-training on this dataset improve downstream classification performance for biomarker scoring?
- RQ5What are the key challenges in developing stain-independent models using real-world diagnostic WSIs?
Key findings
- The ACROBAT dataset contains 4,212 whole-slide images from 1,153 patients, with matched H&E and IHC stains, making it the largest publicly available multi-stain WSI dataset in breast cancer pathology.
- The dataset includes 37,000 manually annotated landmark pairs from 13 annotators, enabling high-precision benchmarking of image registration algorithms.
- The dataset supports diverse computational pathology applications, including stain-guided learning, virtual staining, and artifact detection.
- The data were collected from routine diagnostic workflows, ensuring high clinical relevance and real-world variability.
- The dataset is publicly available via arXiv and the ACROBAT challenge website, with performance metrics accessible for algorithm evaluation.
- The dataset enables the development of stain-independent models by providing consistent anatomical correspondence across staining modalities.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.