[论文解读] A comparative study of artificial intelligence and human doctors for the purpose of triage and diagnosis
本研究在现实情境 vignette 对照下,前瞻性验证 AI 分诊与诊断系统与人类医生的表现,结果显示 AI 的性能可与医生相当,且总体上给出的分诊建议更安全。
Online symptom checkers have significant potential to improve patient care, however their reliability and accuracy remain variable. We hypothesised that an artificial intelligence (AI) powered triage and diagnostic system would compare favourably with human doctors with respect to triage and diagnostic accuracy. We performed a prospective validation study of the accuracy and safety of an AI powered triage and diagnostic system. Identical cases were evaluated by both an AI system and human doctors. Differential diagnoses and triage outcomes were evaluated by an independent judge, who was blinded from knowing the source (AI system or human doctor) of the outcomes. Independently of these cases, vignettes from publicly available resources were also assessed to provide a benchmark to previous studies and the diagnostic component of the MRCGP exam. Overall we found that the Babylon AI powered Triage and Diagnostic System was able to identify the condition modelled by a clinical vignette with accuracy comparable to human doctors (in terms of precision and recall). In addition, we found that the triage advice recommended by the AI System was, on average, safer than that of human doctors, when compared to the ranges of acceptable triage provided by independent expert judges, with only a minimal reduction in appropriateness.
研究动机与目标
- 评估由 AI 支持的分诊与诊断系统(Babylon)相对于人类医生的诊断准确性。
- 评估 AI 驅动的分诊建议的安全性与适宜性。
- 通过半自然情境的 OSCE 设计,检验信息收集和病史询问能力。
- 将 AI 表现在公开可得的病例情境和既定考试材料进行基准比较。
提出的方法
- 采用半自然的角色扮演、模拟会诊,采用 OSCE 格式。
- 将 AI 系统输出与独立盲评评审和多位医生进行比较。
- 使用召回率、精确度和 F1 指标评估鉴别诊断和分诊行动。
- 纳入专家对鉴别诊断质量和分诊安全性的定性评分。
- 通过调整内部阈值以模拟医生风格行为来测试 AI 的敏感性。
实验结果
研究问题
- RQ1AI 驱动的分诊与诊断系统是否能以与人类医生相当的准确率(精确度与召回率)识别情境中的疾病?
- RQ2在独立评审评估阈值下,AI 生成的分诊建议是否与人类医生提供的同样安全或更安全?
- RQ3AI 的表现如何与专家评分的鉴别诊断质量以及既定考试基准相比?
- RQ4调整内部阈值对 AI 系统的召回率与精确度相对于医生的影响是什么?
- RQ5AI 输出是否能推广到公开可用的情境基准(Semigran 2015、MRCGP AKT/CSA)?
主要发现
- AI 系统在各情境中的召回和精确度与医生相当(Babylon AI 召回 80.0%,精确度 44.4%,F1 57.1%)。
- 平均医生召回率:83.9%,精确度:43.6%,F1:57.0%,覆盖七名医生。
- AI 分诊安全性 (97.0%) 超过医生(平均 93.1%),但适宜性相似或略低(AI 90.0% vs 医生 90.5%)。
- 专家评审认为 AI 的鉴别诊断质量可与医生相当(在不同评审组中为 83.0%、83.0%–83.0%?;普通科医生面板结果各异,AI 有时因评估者不同而被评为更低)。
- AI 在 Semigran 2015 情境中的表现:AI 的 top-1 召回率 70.0%,top-3 召回率 96.7%,对比医生的 75.3% 和 90.3%。
- AKT/CSA 基准显示 AI top-3 包含建模疾病的比例为 86.7%(AKT)和 75.0%(CSA)。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。