[논문 리뷰] Performance evaluation of predictive AI models to support medical decisions: Overview and guidance
이 논문은 의학에서 이진 결과 예측 AI에 대해 32개의 성능 지표를 검토하고, 구별력(discrimination), 보정(calibration), 전반(overall), 분류(classification), 임상 활용(clinical utility)을 평가하며, 적절한 지표 선택 및 시각화에 관한 지침을 제공합니다.
A myriad of measures to illustrate performance of predictive artificial intelligence (AI) models have been proposed in the literature. Selecting appropriate performance measures is essential for predictive AI models that are developed to be used in medical practice, because poorly performing models may harm patients and lead to increased costs. We aim to assess the merits of classic and contemporary performance measures when validating predictive AI models for use in medical practice. We focus on models with a binary outcome. We discuss 32 performance measures covering five performance domains (discrimination, calibration, overall, classification, and clinical utility) along with accompanying graphical assessments. The first four domains cover statistical performance, the fifth domain covers decision-analytic performance. We explain why two key characteristics are important when selecting which performance measures to assess: (1) whether the measure's expected value is optimized when it is calculated using the correct probabilities (i.e., a "proper" measure), and (2) whether they reflect either purely statistical performance or decision-analytic performance by properly considering misclassification costs. Seventeen measures exhibit both characteristics, fourteen measures exhibited one characteristic, and one measure possessed neither characteristic (the F1 measure). All classification measures (such as classification accuracy and F1) are improper for clinically relevant decision thresholds other than 0.5 or the prevalence. We recommend the following measures and plots as essential to report: AUROC, calibration plot, a clinical utility measure such as net benefit with decision curve analysis, and a plot with probability distributions per outcome category.
연구 동기 및 목표
- 이진 결과를 갖는 의학 의사결정 지원 AI에 대한 고전적 및 현대적 성능 지표의 장점을 평가한다.
- 지표의 통계적 성능과 의사결정 분석적 성능을 구분한다.
- 적절한(proper) 지표와 의사결정 비용이나 확률을 반영하지 못하는 지표를 식별한다.
- 필수 성능 그래프와 지표를 보고하기 위한 실용적 권고를 제공한다.
제안 방법
- 32개의 성능 지표를 구별력(discrimination), 보정(calibration), 전반(overall), 분류(classification), 임상 활용(clinical utility)의 다섯 영역으로 검토하고 분류한다.
- 각 지표가 적절한지 평가한다(확률이 정확할 때 기대값을 최적화하는지).
- 오판 비용을 고려하여 지표가 통계적 성능과 의사결정 분석적 성능을 반영하는지 평가한다.
- 0.5를 넘는 임계값이나 유병률 외의 일반적으로 임상적으로 관련된 결정 임계값들에 대해 분석한다.
- AUROC, 보정 플롯, 의사결정 곡선 분석이 포함된 순이익(net benefit), 확률 분포 플롯과 같은 필수 보고 항목을 권고한다.
실험 결과
연구 질문
- RQ1의료 현장에서 예측 AI 모델에 대해 어떤 성능 지표가 적절한가?
- RQ2지표들이 순전히 통계적 성능을反영하는가, 아니면 의사결정 분석적 성능을 반영하는가?
- RQ3임상 의사결정을 가장 잘 지원하는 지표와 시각화의 조합은 무엇인가?
주요 결과
- 32개 지표 중 17개가 적절하며 통계적 및/또는 의사결정 분석적 특성을 모두 반영한다.
- 14개 지표는 하나의 바람직한 특성을 보이고, 하나의 지표(F1)는 두 특성을 모두 가지지 않는다.
- 모든 분류 지표(예: 정확도, F1)는 0.5 또는 유병률 외의 임상적으로 관련된 의사결정 임계값에 대해 부적절하다.
- 저자들은 AUROC, 보정 플롯, 순편익(net benefit)과 의사결정 곡선 분석이 포함된 임상 활용 지표, 그리고 결과별 확률 분포 플롯을 보고할 것을 권고한다.
- 본 연구는 의료 현장에서 해와 비용 상승을 피하기 위해 지표를 선택하고 관련 그래픽 평가를 동반하는 지침을 제공한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.