Skip to main content
QUICK REVIEW

[论文解读] Performance evaluation of predictive AI models to support medical decisions: Overview and guidance

Ben Van Calster, Gary S. Collins|arXiv (Cornell University)|Dec 13, 2024
Health Systems, Economic Evaluations, Quality of Life被引用 13
一句话总结

本论文评估了医学领域二元输出预测AI的32个性能度量,评估辨别力、校准、总体、分类和临床实用性,并就适当的度量选择与可视化提供指南。

ABSTRACT

A myriad of measures to illustrate performance of predictive artificial intelligence (AI) models have been proposed in the literature. Selecting appropriate performance measures is essential for predictive AI models that are developed to be used in medical practice, because poorly performing models may harm patients and lead to increased costs. We aim to assess the merits of classic and contemporary performance measures when validating predictive AI models for use in medical practice. We focus on models with a binary outcome. We discuss 32 performance measures covering five performance domains (discrimination, calibration, overall, classification, and clinical utility) along with accompanying graphical assessments. The first four domains cover statistical performance, the fifth domain covers decision-analytic performance. We explain why two key characteristics are important when selecting which performance measures to assess: (1) whether the measure's expected value is optimized when it is calculated using the correct probabilities (i.e., a "proper" measure), and (2) whether they reflect either purely statistical performance or decision-analytic performance by properly considering misclassification costs. Seventeen measures exhibit both characteristics, fourteen measures exhibited one characteristic, and one measure possessed neither characteristic (the F1 measure). All classification measures (such as classification accuracy and F1) are improper for clinically relevant decision thresholds other than 0.5 or the prevalence. We recommend the following measures and plots as essential to report: AUROC, calibration plot, a clinical utility measure such as net benefit with decision curve analysis, and a plot with probability distributions per outcome category.

研究动机与目标

  • 评估用于二元结果的医疗决策支持AI的经典与现代性能度量的优点。
  • 区分统计性能与决策分析性能在度量中的体现。
  • 识别恰当的(proper)度量以及那些未能反映决策成本或概率的度量。
  • 为报告关键性能图和度量提供实用建议。

提出的方法

  • 将32个性能度量归类并分为五个领域:辨别、校准、总体、分类和临床实用性。
  • 评估每个度量是否为proper(在概率正确时优化期望值)。
  • 通过考虑误分类成本来评估度量是否反映统计性能还是决策分析性能。
  • 分析超出0.5或流行度的临床相关决策阈值的常见度量。
  • 建议必须的报告项,包括AUROC、校准图、如净收益与决策曲线分析等临床效用度量,以及按结果的概率分布图。

实验结果

研究问题

  • RQ1哪些度量对医疗实践中的预测AI模型是恰当的?
  • RQ2度量是仅反映统计性能,还是反映决策分析性能?
  • RQ3哪些度量与图表的组合最能支持临床决策?

主要发现

  • 32个度量中有17个是恰当的,且在统计和/或决策分析属性方面都具备。
  • 有14个度量只呈现单一的理想特征,且有一个度量(F1)既无此特征也无此特征。
  • 所有分类度量(如准确度、F1)在临床相关的决策阈值(除0.5或流行度外)都是不恰当的。
  • 作者建议报告AUROC、校准图、如净收益与决策曲线分析等临床效用度量,以及按结果的概率分布图。
  • 本研究为在医疗场景中选择度量和伴随图形评估提供了避免伤害和成本上升的指引。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。