[Paper Review] Mitosis domain generalization in histopathology images -- The MIDOG challenge
This paper presents the MIDOG 2021 challenge, which evaluates domain generalization methods for mitosis detection in histopathology images across multiple whole slide scanners. The winning approach achieved an F₁ score of 0.748 (CI95: 0.704–0.781), demonstrating state-of-the-art performance on unseen scanners, with ensembling, test-time augmentation, and auxiliary segmentation tasks emerging as key success factors.
The density of mitotic figures within tumor tissue is known to be highly correlated with tumor proliferation and thus is an important marker in tumor grading. Recognition of mitotic figures by pathologists is known to be subject to a strong inter-rater bias, which limits the prognostic value. State-of-the-art deep learning methods can support the expert in this assessment but are known to strongly deteriorate when applied in a different clinical environment than was used for training. One decisive component in the underlying domain shift has been identified as the variability caused by using different whole slide scanners. The goal of the MICCAI MIDOG 2021 challenge has been to propose and evaluate methods that counter this domain shift and derive scanner-agnostic mitosis detection algorithms. The challenge used a training set of 200 cases, split across four scanning systems. As a test set, an additional 100 cases split across four scanning systems, including two previously unseen scanners, were given. The best approaches performed on an expert level, with the winning algorithm yielding an F_1 score of 0.748 (CI95: 0.704-0.781). In this paper, we evaluate and compare the approaches that were submitted to the challenge and identify methodological factors contributing to better performance.
Motivation & Objective
- To address domain shift in mitosis detection caused by whole slide scanner variability in digital histopathology.
- To develop scanner-agnostic deep learning models that generalize across different imaging devices.
- To evaluate methods on a benchmark with four scanners, including two unseen in training.
- To identify methodological factors enabling robust performance across diverse scanner domains.
- To advance clinical applicability of AI in tumor grading by reducing inter-rater variability and scanner-dependent performance degradation.
Proposed method
- The challenge used a training set of 200 cases from four whole slide scanners and a test set of 100 cases, including two previously unseen scanners.
- Participants applied deep learning models for mitosis detection, with many employing test-time augmentation and ensembling to improve robustness.
- Top-performing methods incorporated auxiliary segmentation tasks to enhance feature learning and generalization.
- Models were evaluated using F₁ score across all scanners, with special focus on performance on unknown scanners.
- The evaluation framework ensured controlled staining conditions to isolate scanner-induced domain shift as the primary challenge.
- Results were analyzed to identify trends in model architecture, data augmentation, and training strategies influencing domain generalization.
Experimental results
Research questions
- RQ1Can deep learning models achieve robust mitosis detection performance across multiple whole slide scanners, including those not seen during training?
- RQ2What methodological components—such as ensembling, test-time augmentation, or auxiliary tasks—most significantly improve domain generalization in mitosis detection?
- RQ3How does domain shift from scanner differences compare to other sources of variation, such as staining or tissue type, in affecting model performance?
- RQ4To what extent can model performance on region-of-interest (ROI) patches generalize to full whole slide images with higher tissue complexity?
- RQ5What role do model architecture size and design play in achieving generalization across diverse scanner domains?
Key findings
- The best-performing method achieved an F₁ score of 0.748 (CI95: 0.704–0.781), matching the performance level of expert pathologists on the same task.
- All three top-performing approaches used auxiliary segmentation tasks, suggesting a potential benefit for domain generalization, though causality could not be confirmed.
- Five out of the seven top-performing methods used ensembling or test-time augmentation, indicating these strategies as strong contributors to robustness.
- Models showed significantly weaker performance on scanners D and F—especially when labels were not provided—suggesting unaccounted domain shift or image quality issues.
- Larger classification networks were associated with better performance, indicating architectural capacity may support generalization.
- The challenge demonstrated that domain shift from whole slide scanners is a major but surmountable barrier, with top methods achieving state-of-the-art performance on unseen scanners.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.