Skip to main content
QUICK REVIEW

[论文解读] Common Limitations of Image Processing Metrics: A Picture Story

Annika Reinke, Minu D. Tizabi|arXiv (Cornell University)|Apr 12, 2021
Medical Image Segmentation Techniques参考文献 17被引用 87
一句话总结

一个活跃的、以 Delphi 为驱动的概览,详细说明在图像级分类、语义/实例分割和目标检测等方面常见图像处理指标的实际陷阱与局限,并给出上下文感知的指标选择指南。

ABSTRACT

While the importance of automatic image analysis is continuously increasing, recent meta-research revealed major flaws with respect to algorithm validation. Performance metrics are particularly key for meaningful, objective, and transparent performance assessment and validation of the used automatic algorithms, but relatively little attention has been given to the practical pitfalls when using specific metrics for a given image analysis task. These are typically related to (1) the disregard of inherent metric properties, such as the behaviour in the presence of class imbalance or small target structures, (2) the disregard of inherent data set properties, such as the non-independence of the test cases, and (3) the disregard of the actual biomedical domain interest that the metrics should reflect. This living dynamically document has the purpose to illustrate important limitations of performance metrics commonly applied in the field of image analysis. In this context, it focuses on biomedical image analysis problems that can be phrased as image-level classification, semantic segmentation, instance segmentation, or object detection task. The current version is based on a Delphi process on metrics conducted by an international consortium of image analysis experts from more than 60 institutions worldwide.

研究动机与目标

  • 突出指标属性、数据集特征及生物医学领域相关性如何影响医学影像分析中的验证结果。
  • 总结图像级分类、分割和目标检测指标的常见陷阱。
  • 提供面向问题与情境的指标选择指南,以提高可重复性和有效性。

提出的方法

  • 基于混淆矩阵概念(真阳性 TP、假阳性 FP、真阴性 TN、假阴性 FN)回顾并对图像分析中使用的核心指标(计数、多阈值、基于距离的指标)进行分类。
  • 描述指标族之间的关系以及它们如何映射到图像级、对象级和像素级问题。
  • 综合来自>60 institutions 的国际专家联盟进行的 Delphi 过程的见解,以识别陷阱。
  • 提出面向问题类别的具体陷阱(类别-指标不匹配、类别不平衡、多类别问题)以及跨主题的陷阱(聚合、可视化等)。
  • 提出与生物医学相关性和验证目标相一致的选取指标的实用指南。

实验结果

研究问题

  • RQ1在不同问题类别(分类、分割、检测)下,常用图像分析指标的主要局限性和陷阱是什么?
  • RQ2指标属性和数据集特征如何相互作用,影响生物医学影像分析中的验证结果?
  • RQ3可以提供哪些面向问题和情境的指标选择指南,以实现稳健、透明的评估?

主要发现

  • 指标受类别-指标不匹配和数据集属性的深刻影响,导致偏差或误导性的验证结果。
  • 大量与指标相关的陷阱源于将不合适的指标应用于错误的问题类别(例如语义分割与实例分割之分,或图像级与对象级任务之分)。
  • 逐类评估和多类框架可以帮助揭示按类别的性能并避免聚合偏差。
  • 跨主题的陷阱包括无信息的可视化、无效的算法输出以及聚合问题,这些会掩盖真实性能。
  • 本文汇总了来自 BIAs 挑战、MICCAI 挑战和 MONAI 基准测试的指南与工具,以促进情境感知的指标选择。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。