[Paper Review] Towards Efficient COVID-19 CT Annotation: A Benchmark for Lung and Infection Segmentation
This paper introduces a publicly available, large-scale 3D COVID-19 CT dataset with over 1,800 annotated slices across 20 cases, along with three standardized benchmarks for lung and infection segmentation. It enables fair comparison of deep learning methods through unified data splits, evaluation metrics, and 40+ pre-trained models, advancing annotation-efficient segmentation under data-scarce conditions.
Accurate segmentation of lung and infection in COVID-19 CT scans plays an important role in the quantitative management of patients. Most of the existing studies are based on large and private annotated datasets that are impractical to obtain from a single institution, especially when radiologists are busy fighting the coronavirus disease. Furthermore, it is hard to compare current COVID-19 CT segmentation methods as they are developed on different datasets, trained in different settings, and evaluated with different metrics. In this paper, we created a COVID-19 3D CT dataset with 20 cases that contains 1800+ annotated slices and made it publicly available. To promote the development of annotation-efficient deep learning methods, we built three benchmarks for lung and infection segmentation that contain current main research interests, e.g., few-shot learning, domain generalization, and knowledge transfer. For a fair comparison among different segmentation methods, we also provide unified training, validation and testing dataset splits, and evaluation metrics and corresponding code. In addition, we provided more than 40 pre-trained baseline models for the benchmarks, which not only serve as out-of-the-box segmentation tools but also save computational time for researchers who are interested in COVID-19 lung and infection segmentation. To the best of our knowledge, this work presents the largest public annotated COVID-19 CT volume dataset, the first segmentation benchmark, and the most pre-trained models up to now. We hope these resources (\url{this https URL}) could advance the development of deep learning methods for COVID-19 CT segmentation with limited data.
Motivation & Objective
- To address the lack of publicly available, large-scale, and well-annotated COVID-19 CT datasets for research.
- To overcome the challenge of comparing segmentation methods due to inconsistent datasets, training protocols, and evaluation metrics.
- To enable annotation-efficient deep learning by establishing benchmarks for few-shot learning, domain generalization, and knowledge transfer.
- To provide a standardized platform with unified training, validation, and testing splits for reproducible research.
- To accelerate method development by releasing over 40 pre-trained models for immediate use and reduced computational cost.
Proposed method
- Constructed a public 3D COVID-19 CT dataset comprising 20 cases with more than 1,800 annotated slices for lung and infection regions.
- Designed three segmentation benchmarks targeting key research challenges: few-shot learning, domain generalization, and knowledge transfer.
- Established standardized, publicly shared dataset splits (train/val/test) to ensure fair and reproducible model evaluation.
- Defined consistent evaluation metrics and released corresponding code to unify performance comparison across methods.
- Trained and released over 40 pre-trained deep learning models for the benchmarks to serve as strong baselines and reduce training time.
- Focused on practical utility by aligning benchmarks with real-world constraints such as limited annotations and data scarcity.
Experimental results
Research questions
- RQ1How can we establish a standardized benchmark for evaluating lung and infection segmentation in COVID-19 CT scans across diverse deep learning methods?
- RQ2What is the performance of state-of-the-art models under annotation-scarce settings such as few-shot learning and domain generalization?
- RQ3To what extent do pre-trained models from this benchmark improve downstream segmentation performance with minimal fine-tuning?
- RQ4How does the unified evaluation protocol improve reproducibility and comparability of segmentation methods compared to prior work?
- RQ5Can knowledge transfer from pre-trained models significantly reduce the need for extensive annotation in clinical CT segmentation tasks?
Key findings
- This work presents the largest public annotated 3D COVID-19 CT dataset to date, with 20 cases and over 1,800 annotated slices.
- The study introduces the first standardized benchmark for lung and infection segmentation in COVID-19 CT with unified data splits and evaluation protocols.
- More than 40 pre-trained models are released, providing strong baselines and reducing computational overhead for new research.
- The unified evaluation framework enables fair and reproducible comparison of segmentation models across different methods and settings.
- The benchmark supports key research directions such as few-shot learning, domain generalization, and knowledge transfer in medical image segmentation.
- The public availability of the dataset and tools is expected to accelerate the development of data-efficient deep learning methods for clinical CT analysis.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.