Skip to main content
QUICK REVIEW

[Paper Review] PolypGen: A multi-center polyp detection and segmentation dataset for generalisability assessment.

Sharib Ali, Debesh Jha|arXiv (Cornell University)|Jun 8, 2021
Colorectal Cancer Screening and Detection32 references9 citations
TL;DR

PolypGen is a large-scale, multi-center dataset of 3,446 precisely annotated polyp instances across 300+ patients, curated for rigorous evaluation of polyp detection and segmentation models. It enables generalisability assessment through diverse endoscopic data from six centers, verified by six senior gastroenterologists, and was developed as part of the Endocv2021 challenge.

ABSTRACT

Polyps in the colon are widely known as cancer precursors identified by colonoscopy either related to diagnostic work-up for symptoms, colorectal cancer screening or systematic surveillance of certain diseases. Whilst most polyps are benign, the number, size and the surface structure of the polyp are tightly linked to the risk of colon cancer. There exists a high missed detection rate and incomplete removal of colon polyps due to the variable nature, difficulties to delineate the abnormality, high recurrence rates and the anatomical topography of the colon. In the past, several methods have been built to automate polyp detection and segmentation. However, the key issue of most methods is that they have not been tested rigorously on a large multi-center purpose-built dataset. Thus, these methods may not generalise to different population datasets as they overfit to a specific population and endoscopic surveillance. To this extent, we have curated a dataset from 6 different centers incorporating more than 300 patients. The dataset includes both single frame and sequence data with 3446 annotated polyp labels with precise delineation of polyp boundaries verified by six senior gastroenterologists. To our knowledge, this is the most comprehensive detection and pixel-level segmentation dataset curated by a team of computational scientists and expert gastroenterologists. This dataset has been originated as the part of the Endocv2021 challenge aimed at addressing generalisability in polyp detection and segmentation. In this paper, we provide comprehensive insight into data construction and annotation strategies, annotation quality assurance and technical validation for our extended EndoCV2021 dataset which we refer to as PolypGen.

Motivation & Objective

  • To address the lack of large-scale, multi-center datasets for evaluating the generalisability of polyp detection and segmentation models.
  • To reduce overfitting to single-center or single-population data by creating a diverse, multi-institutional dataset with standardized annotation.
  • To support robust benchmarking of deep learning models in real-world clinical settings by incorporating both single-frame and sequence data.
  • To ensure high annotation quality through consensus labeling by six expert gastroenterologists.
  • To provide a foundation for future research in colonoscopy image analysis by offering a comprehensive, validated dataset for generalisability assessment.

Proposed method

  • Curated endoscopic images and video sequences from six independent medical centers to ensure diversity in patient demographics, endoscope types, and imaging protocols.
  • Collected 3,446 polyp instances annotated with precise pixel-level segmentation masks by six senior gastroenterologists to ensure high-quality ground truth.
  • Implemented a multi-reader consensus annotation strategy to enhance reliability and reduce inter-observer variability in polyp boundary delineation.
  • Organized the dataset into single-frame and video sequence subsets to support both classification and temporal modeling tasks.
  • Validated data quality through technical checks, including spatial consistency, annotation completeness, and outlier detection.
  • Released the dataset as an extension of the Endocv2021 challenge to promote community-wide evaluation and benchmarking of polyp detection and segmentation models.

Experimental results

Research questions

  • RQ1To what extent does model performance degrade when evaluated on a multi-center dataset compared to single-center benchmarks?
  • RQ2How does the inclusion of diverse endoscopic data from multiple institutions affect the generalisability of polyp detection and segmentation models?
  • RQ3What is the impact of expert consensus annotation on the reliability and reproducibility of polyp segmentation in clinical imaging?
  • RQ4Can a multi-center dataset like PolypGen serve as a robust benchmark for evaluating model generalisability across different populations and imaging conditions?
  • RQ5How does the diversity of anatomical locations, polyp morphology, and imaging quality in PolypGen influence model robustness?

Key findings

  • PolypGen contains 3,446 polyp instances with pixel-level segmentation masks, making it the most comprehensive dataset for polyp detection and segmentation to date.
  • The dataset spans six medical centers, ensuring diversity in patient populations, endoscopic equipment, and imaging protocols to support generalisability testing.
  • All annotations were verified by six senior gastroenterologists, ensuring high inter-rater reliability and clinical relevance.
  • The inclusion of both single-frame and video sequence data enables evaluation of models across different temporal and spatial contexts.
  • The dataset was developed as part of the Endocv2021 challenge, establishing a standardized benchmark for evaluating model performance across diverse clinical settings.
  • PolypGen provides a foundation for future research by enabling rigorous assessment of model generalisability beyond single-center or single-population performance.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.