Skip to main content
QUICK REVIEW

[Paper Review] Performance evaluation of predictive AI models to support medical decisions: Overview and guidance

Ben Van Calster, Gary S. Collins|arXiv (Cornell University)|Dec 13, 2024
Health Systems, Economic Evaluations, Quality of Life13 citations
TL;DR

This paper reviews 32 performance measures for binary-p outcome predictive AI in medicine, evaluating discrimination, calibration, overall, classification, and clinical utility, and provides guidance on proper measure selection and visualization.

ABSTRACT

A myriad of measures to illustrate performance of predictive artificial intelligence (AI) models have been proposed in the literature. Selecting appropriate performance measures is essential for predictive AI models that are developed to be used in medical practice, because poorly performing models may harm patients and lead to increased costs. We aim to assess the merits of classic and contemporary performance measures when validating predictive AI models for use in medical practice. We focus on models with a binary outcome. We discuss 32 performance measures covering five performance domains (discrimination, calibration, overall, classification, and clinical utility) along with accompanying graphical assessments. The first four domains cover statistical performance, the fifth domain covers decision-analytic performance. We explain why two key characteristics are important when selecting which performance measures to assess: (1) whether the measure's expected value is optimized when it is calculated using the correct probabilities (i.e., a "proper" measure), and (2) whether they reflect either purely statistical performance or decision-analytic performance by properly considering misclassification costs. Seventeen measures exhibit both characteristics, fourteen measures exhibited one characteristic, and one measure possessed neither characteristic (the F1 measure). All classification measures (such as classification accuracy and F1) are improper for clinically relevant decision thresholds other than 0.5 or the prevalence. We recommend the following measures and plots as essential to report: AUROC, calibration plot, a clinical utility measure such as net benefit with decision curve analysis, and a plot with probability distributions per outcome category.

Motivation & Objective

  • Assess the merits of classic and contemporary performance measures for medical decision-support AI with binary outcomes.
  • Distinguish between statistical and decision-analytic performance in measures.
  • Identify proper (proper) measures and those that fail to reflect decision costs or probabilities.
  • Provide practical recommendations for reporting essential performance plots and measures.

Proposed method

  • Review and classify 32 performance measures into five domains: discrimination, calibration, overall, classification, and clinical utility.
  • Evaluate whether each measure is proper (optimizes expected value when probabilities are correct).
  • Assess whether measures reflect statistical versus decision-analytic performance by accounting for misclassification costs.
  • Analyze measures for common clinically relevant decision thresholds beyond 0.5 or prevalence.
  • Recommend essential reporting items including AUROC, calibration plots, net benefit with decision curve analysis, and distribution plots.

Experimental results

Research questions

  • RQ1Which performance measures are proper for predictive AI models in medical practice?
  • RQ2Do measures reflect purely statistical performance or decision-analytic performance?
  • RQ3What combination of measures and plots best supports clinical decision-making?

Key findings

  • Seventeen of the 32 measures are proper and reflect both statistical and/or decision-analytic properties.
  • Fourteen measures show a single desirable characteristic, and one measure (F1) has neither characteristic.
  • All classification measures (e.g., accuracy, F1) are improper for clinically relevant decision thresholds other than 0.5 or the prevalence.
  • The authors recommend reporting AUROC, calibration plots, a clinical utility measure such as net benefit with decision curve analysis, and probability-distribution plots by outcome.
  • The work provides guidance on selecting measures and accompanying graphical assessments to avoid harm and cost escalation in medical settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.