Skip to main content
QUICK REVIEW

[Paper Review] Common Limitations of Image Processing Metrics: A Picture Story

Annika Reinke, Minu D. Tizabi|arXiv (Cornell University)|Apr 12, 2021
Medical Image Segmentation Techniques17 references87 citations
TL;DR

A living, Delphi-driven overview detailing practical pitfalls and limitations of common image processing metrics across image-level classification, semantic/instance segmentation, and object detection, with guidelines for context-aware metric selection.

ABSTRACT

While the importance of automatic image analysis is continuously increasing, recent meta-research revealed major flaws with respect to algorithm validation. Performance metrics are particularly key for meaningful, objective, and transparent performance assessment and validation of the used automatic algorithms, but relatively little attention has been given to the practical pitfalls when using specific metrics for a given image analysis task. These are typically related to (1) the disregard of inherent metric properties, such as the behaviour in the presence of class imbalance or small target structures, (2) the disregard of inherent data set properties, such as the non-independence of the test cases, and (3) the disregard of the actual biomedical domain interest that the metrics should reflect. This living dynamically document has the purpose to illustrate important limitations of performance metrics commonly applied in the field of image analysis. In this context, it focuses on biomedical image analysis problems that can be phrased as image-level classification, semantic segmentation, instance segmentation, or object detection task. The current version is based on a Delphi process on metrics conducted by an international consortium of image analysis experts from more than 60 institutions worldwide.

Motivation & Objective

  • Highlight how metric properties, dataset characteristics, and biomedical domain relevance affect validation outcomes in medical image analysis.
  • Summarize common pitfalls for image-level classification, segmentation, and object detection metrics.
  • Provide guidance for problem- and context-aware metric selection to improve reproducibility and validity.

Proposed method

  • Review and categorize core metrics used in image analysis (counting, multi-threshold, and distance-based) based on confusion matrix concepts (TP, FP, TN, FN).
  • Describe relationships among metric families and how they map to image-, object-, and pixel-level problems.
  • Synthesize insights from a Delphi process conducted by an international consortium of experts from >60 institutions to identify pitfalls.
  • Present problem-category specific pitfalls (category-metric mismatch, class imbalance, multi-class issues) and cross-topic pitfalls (aggregation, visualization, etc.).
  • Frame practical guidelines for choosing metrics aligned with biomedical relevance and validation goals.

Experimental results

Research questions

  • RQ1What are the major limitations and pitfalls of commonly used image analysis metrics across different problem categories (classification, segmentation, detection)?
  • RQ2How do metric properties and dataset characteristics interact to affect validation outcomes in biomedical image analysis?
  • RQ3What guidelines can be provided to select metrics in a problem- and context-aware manner for robust, transparent evaluation?

Key findings

  • Metrics are profoundly influenced by category-metric mismatch and dataset properties, leading to biased or misleading validation outcomes.
  • A large portion of metric-related pitfalls arises from applying inappropriate metrics to the wrong problem category (e.g., semantic vs. instance segmentation, or image-level vs. object-level tasks).
  • Per-class evaluation and multi-class framing can help reveal class-wise performance and avoid aggregation bias.
  • Cross-topic pitfalls include uninformative visualizations, invalid algorithm outputs, and aggregation issues that obscure true performance.
  • The document consolidates guidelines and tools from BIAs challenges, MICCAI challenges, and MONAI benchmarking to promote context-aware metric selection.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.