[Paper Review] Beyond Benchmarks of IUGC: Rethinking Requirements of Deep Learning Methods for Intrapartum Ultrasound Biometry from Fetal Ultrasound Videos
This paper analyzes the Intrapartum Ultrasound Grand Challenge (IUGC) 2024, reviewing multi-task deep learning approaches for standard plane classification, FH-PS segmentation, and biometry from intrapartum ultrasound videos, and discusses dataset, methods, and challenges.
A substantial proportion (45\%) of maternal deaths, neonatal deaths, and stillbirths occur during the intrapartum phase, with a particularly high burden in low- and middle-income countries. Intrapartum biometry plays a critical role in monitoring labor progression; however, the routine use of ultrasound in resource-limited settings is hindered by a shortage of trained sonographers. To address this challenge, the Intrapartum Ultrasound Grand Challenge (IUGC), co-hosted with MICCAI 2024, was launched. The IUGC introduces a clinically oriented multi-task automatic measurement framework that integrates standard plane classification, fetal head-pubic symphysis segmentation, and biometry, enabling algorithms to exploit complementary task information for more accurate estimation. Furthermore, the challenge releases the largest multi-center intrapartum ultrasound video dataset to date, comprising 774 videos (68,106 frames) collected from three hospitals, providing a robust foundation for model training and evaluation. In this study, we present a comprehensive overview of the challenge design, review the submissions from eight participating teams, and analyze their methods from five perspectives: preprocessing, data augmentation, learning strategy, model architecture, and post-processing. In addition, we perform a systematic analysis of the benchmark results to identify key bottlenecks, explore potential solutions, and highlight open challenges for future research. Although encouraging performance has been achieved, our findings indicate that the field remains at an early stage, and further in-depth investigation is required before large-scale clinical deployment. All benchmark solutions and the complete dataset have been publicly released to facilitate reproducible research and promote continued advances in automatic intrapartum ultrasound biometry.
Motivation & Objective
- Motivate automated, end-to-end measurement of intrapartum ultrasound parameters to support labor management and reduce adverse outcomes.
- Provide a standardized multi-center intrapartum video benchmark for tasks including standard plane classification, FH-PS segmentation, and biometry (AoP and HSD).
- Review submitted methods across preprocessing, data augmentation, learning strategy, architecture, and post-processing to identify key bottlenecks and opportunities for improvement.
- Facilitate open scientific exchange by releasing datasets and code to drive future research in intrapartum ultrasound biometry.
Proposed method
- Describe a multitask framework that integrates standard plane classification, FH-PS segmentation, and biometry from ultrasound videos.
- Analyze challenge submissions and organize evaluation using the Biomedical Image Analysis ChallengeS (BIAS) methodology.
- Leverage a large multi-center video dataset (774 videos, 68,106 images) for training, validation, and testing.
- Provide public access to solutions and datasets to enable reproducibility and further development.
- Discuss use of template/landmark-based measurement for AoP and HSD within segmented FH and PS regions.
- Highlight challenges such as intra-class variation, low inter-class differences, ultrasound noise, and anatomical deformation during labor.

Experimental results
Research questions
- RQ1What are effective deep learning strategies for multi-task intrapartum ultrasound analysis (plane classification, FH-PS segmentation, and biometry) from video data?
- RQ2How does a large multi-center intrapartum video dataset influence generalization and benchmark reliability for automatic biometry?
- RQ3What are the main obstacles and potential solutions to achieving clinically usable automatic intrapartum biometry?
- RQ4To what extent can public datasets and open-source baselines accelerate progress in end-to-end intrapartum ultrasound measurement?
Key findings
- The IUGC 2024 dataset comprises 774 videos (68,106 images) from three hospitals, with training/validation/test splits of 434/40/300 videos respectively.
- Eight teams submitted across classification, segmentation, and biometry with seven top-performing methods analyzed alongside a baseline.
- Multitask approaches were evaluated in a real-world, multi-center setting, highlighting the importance of integrating standard plane recognition with FH-PS segmentation for accurate biometry.
- Despite promising results, the study emphasizes that intrapartum ultrasound biometry remains in early stages and further work is required before clinical deployment.
- The authors provide public access to the complete dataset and the top-performing methods to foster continued advancement in automatic intrapartum biometry.
- The paper discusses ongoing clinical and technical challenges, including data heterogeneity, spatio-temporal modeling of ultrasound videos, and error propagation across cascaded tasks.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.