[Paper Review] AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image Segmentation
AMOS provides a large-scale, diverse abdominal CT/MRI dataset with voxel-level annotations for 15 organs, enabling robust multi-organ segmentation and cross-domain evaluations. It benchmarks current methods and demonstrates the dataset’s utility for OOD generalization, cross-modality learning, and transfer learning.
Despite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse clinical scenarios. Constraint by the high cost of collecting and labeling 3D medical data, most of the deep learning models to date are driven by datasets with a limited number of organs of interest or samples, which still limits the power of modern deep models and makes it difficult to provide a fully comprehensive and fair estimate of various methods. To mitigate the limitations, we present AMOS, a large-scale, diverse, clinical dataset for abdominal organ segmentation. AMOS provides 500 CT and 100 MRI scans collected from multi-center, multi-vendor, multi-modality, multi-phase, multi-disease patients, each with voxel-level annotations of 15 abdominal organs, providing challenging examples and test-bed for studying robust segmentation algorithms under diverse targets and scenarios. We further benchmark several state-of-the-art medical segmentation models to evaluate the status of the existing methods on this new challenging dataset. We have made our datasets, benchmark servers, and baselines publicly available, and hope to inspire future research. Information can be found at https://amos22.grand-challenge.org.
Motivation & Objective
- Address the lack of large-scale, diverse abdominal segmentation benchmarks that reflect real-world clinical variability.
- Provide a multi-modality (CT and MRI), multi-center, multi-scanner, multi-phase, multi-disease dataset with dense organ annotations.
- Benchmark state-of-the-art segmentation models on AMOS to assess current limitations and robustness.
- Demonstrate AMOS's versatility for tasks beyond segmentation, including OOD generalization, cross-modality learning, and transfer learning.
Proposed method
- Curate a large, diverse abdominal dataset (CT and MRI) with voxel-level annotations for 15 organs from two clinical centers and multiple scanners.
- Use a semi-automatic annotation workflow: coarse labels from pre-trained segmentors followed by iterative refinement by junior and senior radiologists to ensure quality.
- Define data splits that include in-distribution (ID) and out-of-distribution (OOD) evaluations across scanners to measure domain robustness.
- Benchmark six baselines (CNN, Transformer, and hybrid methods) on AMOS-CT and AMOS-MRI using standard metrics (Dice and NSD) and report model efficiency (parameters and FLOPs).
- Explore extended tasks enabled by AMOS, including cross-modality learning, transfer learning, and privacy-preserving/federated settings.
- Provide dataset, benchmark servers, and baselines publicly to facilitate community research.
Experimental results
Research questions
- RQ1How does state-of-the-art abdominal multi-organ segmentation perform on a large-scale, diverse dataset with 15 organs across CT and MRI modalities?
- RQ2What is the impact of domain shift (different scanners/vendors) on segmentation performance, and can AMOS enable robust OOD generalization studies?
- RQ3Can cross-modality training (CT and MRI) improve segmentation performance for each modality?
- RQ4Do pre-trained representations on AMOS offer transferable benefits to external abdominal segmentation tasks (transfer learning)?
- RQ5What are the data diversity and annotation quality implications for model learning on AMOS?
Key findings
- AMOS comprises 600 scans (500 CT and 100 MRI) with 74,026 annotated slices across 15 abdominal organs, making it the largest and most diverse abdominal segmentation benchmark to date.
- Baseline experiments reveal that simple UNet-like models can outperform several newer architectures on AMOS, and Transformer-based methods do not consistently surpass CNN-based models within this dataset setting.
- There is a notable performance gap between in-distribution (ID) and out-of-distribution (OOD) test data, especially for AMOS-MRI, highlighting domain shift challenges across scanners.
- Cross-modality learning (joint CT+MRI training) yields consistent gains over single-modality training, indicating complementary information between CT and MRI enhances segmentation performance.
- Pre-training on AMOS provides transferable benefits to related abdominal segmentation tasks in several benchmarks, though benefits vary by target domain and modality.
- AMOS enables robust evaluation of generalization, cross-modality learning, and transfer learning, underscoring its potential as a versatile benchmark for real-world clinical scenarios.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.