Skip to main content
QUICK REVIEW

[論文レビュー] Performance evaluation of predictive AI models to support medical decisions: Overview and guidance

Ben Van Calster, Gary S. Collins|arXiv (Cornell University)|Dec 13, 2024
Health Systems, Economic Evaluations, Quality of Life被引用数 13
ひとこと要約

この論文は、医療における二値出力予測AIの32の性能指標をレビューし、識別性、キャリブレーション、全体、分類、臨床有用性を評価し、適切な指標選択と可視化に関するガイダンスを提供する。

ABSTRACT

A myriad of measures to illustrate performance of predictive artificial intelligence (AI) models have been proposed in the literature. Selecting appropriate performance measures is essential for predictive AI models that are developed to be used in medical practice, because poorly performing models may harm patients and lead to increased costs. We aim to assess the merits of classic and contemporary performance measures when validating predictive AI models for use in medical practice. We focus on models with a binary outcome. We discuss 32 performance measures covering five performance domains (discrimination, calibration, overall, classification, and clinical utility) along with accompanying graphical assessments. The first four domains cover statistical performance, the fifth domain covers decision-analytic performance. We explain why two key characteristics are important when selecting which performance measures to assess: (1) whether the measure's expected value is optimized when it is calculated using the correct probabilities (i.e., a "proper" measure), and (2) whether they reflect either purely statistical performance or decision-analytic performance by properly considering misclassification costs. Seventeen measures exhibit both characteristics, fourteen measures exhibited one characteristic, and one measure possessed neither characteristic (the F1 measure). All classification measures (such as classification accuracy and F1) are improper for clinically relevant decision thresholds other than 0.5 or the prevalence. We recommend the following measures and plots as essential to report: AUROC, calibration plot, a clinical utility measure such as net benefit with decision curve analysis, and a plot with probability distributions per outcome category.

研究の動機と目的

  • 二値アウトカムを伴う医療意思決定支援AIの古典的・現代的な性能指標の利点を評価する。
  • 指標における統計的性能と意思決定分析的性能を区別する。
  • 適切な指標(proper)と意思決定コストや確率を反映しない指標を識別する。
  • 必須の性能プロットと指標を報告するための実践的な推奨を提供する。

提案手法

  • 32の性能指標を識別・分類し、識別性・キャリブレーション・全体・分類・臨床有用性の5領域に分類する。
  • 各指標が適切であるか(確率が正しいとき期待値を最適化するか)を評価する。
  • 誤分類コストを考慮して、統計的性能と意思決定分析的性能のどちらを反映しているかを評価する。
  • 0.5や有病率以外の臨床的に関連する一般的な意思決定閾値を分析する。
  • AUROC・キャリブレーションプロット・意思決定曲線分析を含む純利益(net benefit)など、必須の報告項目と分布プロットを推奨する。

実験結果

リサーチクエスチョン

  • RQ1医療実務における予測AIモデルの適切な性能指標はどれか。
  • RQ2指標は純粋に統計的性能を反映するのか、それとも意思決定分析的性能を反映するのか。
  • RQ3臨床意思決定を最もよく支える指標とプロットの組み合わせは何か。

主な発見

  • 32指標のうち17は適切であり、統計的性能および/または意思決定分析的特性の両方を反映している。
  • 14指標は1つの望ましい特性を示し、1つの指標(F1)はどちらの特性も持たない。
  • すべての分類指標(例:accuracy, F1)は、0.5または有病率以外の臨床的に関連する意思決定閾値には適切でない。
  • 著者らはAUROC・キャリブレーションプロット・純利益のような臨床有用性指標(意思決定曲線分析とともに)および結果別の確率分布プロットを報告することを推奨している。
  • 本研究は医療現場での害と費用の増大を回避するための指標選択と付随グラフィカル評価のガイダンスを提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。