Skip to main content
QUICK REVIEW

[Paper Review] Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)

Noel Codella, Veronica Rotemberg|arXiv (Cornell University)|Feb 9, 2019
Cutaneous Melanoma Detection and ManagementMedicine7 references984 citations
TL;DR

This paper summarizes the ISIC 2018 Challenge on Skin Lesion Analysis for melanoma detection, detailing dataset, tasks, evaluation protocols, results, and implications for generalization and regulation.

ABSTRACT

This work summarizes the results of the largest skin image analysis challenge in the world, hosted by the International Skin Imaging Collaboration (ISIC), a global partnership that has organized the world's largest public repository of dermoscopic images of skin. The challenge was hosted in 2018 at the Medical Image Computing and Computer Assisted Intervention (MICCAI) conference in Granada, Spain. The dataset included over 12,500 images across 3 tasks. 900 users registered for data download, 115 submitted to the lesion segmentation task, 25 submitted to the lesion attribute detection task, and 159 submitted to the disease classification task. Novel evaluation protocols were established, including a new test for segmentation algorithm performance, and a test for algorithm ability to generalize. Results show that top segmentation algorithms still fail on over 10% of images on average, and algorithms with equal performance on test data can have different abilities to generalize. This is an important consideration for agencies regulating the growing set of machine learning tools in the healthcare domain, and sets a new standard for future public challenges in healthcare.

Motivation & Objective

  • Present the ISIC 2018 Challenge design and participation metrics.
  • Introduce new evaluation protocols including Thresholded Jaccard and balanced accuracy.
  • Assess generalization by using internal and external test partitions.
  • Analyze results across segmentation, attribute detection, and disease classification tasks.
  • Provide recommendations for future public challenges in healthcare ML.

Proposed method

  • Divide the challenge into three tasks: segmentation, attribute detection, and disease classification.
  • Use Thresholded Jaccard to account for interobserver variability in segmentation.
  • Use balanced accuracy to mitigate prevalence bias in classification.
  • Include internal and external held-out test partitions to evaluate generalization.
  • Provide 4-page manuscripts describing methods and disclose use of in-domain or out-of-domain data.
  • Evaluate segmentation, attribute detection, and classification with task-specific metrics.

Experimental results

Research questions

  • RQ1How do segmentation, attribute detection, and disease classification perform under new evaluation protocols?
  • RQ2Does Thresholded Jaccard better reflect clinical utility than Jaccard in segmentation?
  • RQ3How does balanced accuracy affect ranking and generalization compared to other metrics?
  • RQ4Can algorithms generalize from internal to external data partitions in melanoma detection?
  • RQ5What are the implications of limited attribute detection performance for clinical practice and future challenges?

Key findings

  • Top segmentation submissions reached around 0.80 Thresholded Jaccard but still fail on over 10% of images.
  • Attribute detection performance was low, with best average Jaccard around 0.473 per attribute.
  • Highest disease classification balanced accuracy was 0.885, with notable internal-external generalization gaps.
  • Algorithms often overfit to internal data; generalization varied across methods.
  • Balanced accuracy significantly impacts participant ranking compared with accuracy or AUC.
  • External test data revealed differences in performance not captured by internal test datasets.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.