Skip to main content
QUICK REVIEW

[Paper Review] Overview of the HECKTOR Challenge at MICCAI 2021: Automatic Head and Neck Tumor Segmentation and Outcome Prediction in PET/CT Images

Vincent Andrearczyk, Valentin Oreiller|arXiv (Cornell University)|Jan 11, 2022
Radiomics and Machine Learning in Medical Imaging9 citations
TL;DR

This paper presents the HECKTOR challenge at MICCAI 2021, which evaluated automated segmentation of head and neck tumors and progression-free survival prediction in FDG-PET/CT images using deep learning. The best models achieved a Dice score of 0.7591 for tumor segmentation and C-index values of 0.7196 and 0.6978 for survival prediction with and without ground truth tumor contours, indicating that fully automatic methods can achieve high performance without manual GTV delineation.

ABSTRACT

This paper presents an overview of the second edition of the HEad and neCK TumOR (HECKTOR) challenge, organized as a satellite event of the 24th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2021. The challenge is composed of three tasks related to the automatic analysis of PET/CT images for patients with Head and Neck cancer (H&N), focusing on the oropharynx region. Task 1 is the automatic segmentation of H&N primary Gross Tumor Volume (GTVt) in FDG-PET/CT images. Task 2 is the automatic prediction of Progression Free Survival (PFS) from the same FDG-PET/CT. Finally, Task 3 is the same as Task 2 with ground truth GTVt annotations provided to the participants. The data were collected from six centers for a total of 325 images, split into 224 training and 101 testing cases. The interest in the challenge was highlighted by the important participation with 103 registered teams and 448 result submissions. The best methods obtained a Dice Similarity Coefficient (DSC) of 0.7591 in the first task, and a Concordance index (C-index) of 0.7196 and 0.6978 in Tasks 2 and 3, respectively. In all tasks, simplicity of the approach was found to be key to ensure generalization performance. The comparison of the PFS prediction performance in Tasks 2 and 3 suggests that providing the GTVt contour was not crucial to achieve best results, which indicates that fully automatic methods can be used. This potentially obviates the need for GTVt contouring, opening avenues for reproducible and large scale radiomics studies including thousands potential subjects.

Motivation & Objective

  • To develop and evaluate automated methods for segmenting primary head and neck tumors in FDG-PET/CT images.
  • To predict progression-free survival (PFS) from PET/CT images without relying on manual tumor delineation.
  • To assess whether providing ground truth tumor contours improves PFS prediction performance.
  • To enable large-scale, reproducible radiomics studies by reducing dependence on time-consuming manual segmentation.
  • To evaluate the generalization of deep learning models across diverse imaging protocols from multiple centers.

Proposed method

  • The challenge used a multi-center dataset of 325 FDG-PET/CT scans from six institutions, split into 224 training and 101 testing cases.
  • Participants applied deep learning models, primarily 3D U-Net architectures, for tumor segmentation (Task 1) and survival prediction (Tasks 2 and 3).
  • For survival prediction, models were trained on radiomic features extracted from PET and CT images, with or without prior GTV segmentation.
  • Image preprocessing included standardized reconstruction protocols, including OSEM iterative reconstruction with time-of-flight and resolution modeling.
  • Evaluation metrics included Dice Similarity Coefficient (DSC) for segmentation and Concordance index (C-index) for survival prediction.
  • The challenge used a standardized evaluation platform on AIcrowd to ensure reproducibility and fairness across submissions.

Experimental results

Research questions

  • RQ1Can deep learning models achieve high-performance automatic segmentation of head and neck tumors in multi-center FDG-PET/CT scans?
  • RQ2Does providing ground truth tumor contours significantly improve the accuracy of progression-free survival prediction?
  • RQ3To what extent can fully automatic methods replace manual GTV delineation in radiomics-based outcome prediction?
  • RQ4How generalizable are deep learning models across diverse imaging protocols and scanner types in multi-center medical imaging?
  • RQ5Can automated segmentation and outcome prediction pipelines support large-scale, reproducible radiomics studies?

Key findings

  • The best-performing model achieved a Dice Similarity Coefficient (DSC) of 0.7591 in tumor segmentation, indicating strong agreement with manual contours.
  • For progression-free survival prediction, the best model achieved a C-index of 0.7196 in Task 2 (without GTV annotation), demonstrating high predictive performance.
  • The C-index for Task 3 (with GTV annotation) was 0.6978, suggesting that GTV annotation did not significantly improve performance.
  • The results indicate that fully automatic methods can achieve high performance without manual GTV delineation, reducing dependency on time-consuming manual contouring.
  • Simplicity of model architecture was found to be a key factor in achieving good generalization across diverse imaging centers and protocols.
  • The challenge demonstrated the feasibility of large-scale, reproducible radiomics studies using automated segmentation pipelines.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.