[论文解读] AI-based Clinical Decision Support for Primary Care: A Real-World Study
本研究评估在肯尼亚内罗毕部署的基于LLM的临床决策支持工具(AI Consult),显示实时治疗中临床错误减少、临床医生反馈积极,但患者自述结局无显著差异。
We evaluate the impact of large language model-based clinical decision support in live care. In partnership with Penda Health, a network of primary care clinics in Nairobi, Kenya, we studied AI Consult, a tool that serves as a safety net for clinicians by identifying potential documentation and clinical decision-making errors. AI Consult integrates into clinician workflows, activating only when needed and preserving clinician autonomy. We conducted a quality improvement study, comparing outcomes for 39,849 patient visits performed by clinicians with or without access to AI Consult across 15 clinics. Visits were rated by independent physicians to identify clinical errors. Clinicians with access to AI Consult made relatively fewer errors: 16% fewer diagnostic errors and 13% fewer treatment errors. In absolute terms, the introduction of AI Consult would avert diagnostic errors in 22,000 visits and treatment errors in 29,000 visits annually at Penda alone. In a survey of clinicians with AI Consult, all clinicians said that AI Consult improved the quality of care they delivered, with 75% saying the effect was "substantial". These results required a clinical workflow-aligned AI Consult implementation and active deployment to encourage clinician uptake. We hope this study demonstrates the potential for LLM-based clinical decision support tools to reduce errors in real-world settings and provides a practical framework for advancing responsible adoption.
研究动机与目标
- 评估基于LLM的 CDS 工具是否能减少初级保健中的临床文档和决策错误。
- 评估临床对齐的实施和主动部署如何影响采用率和有效性。
- 在真实世界环境中描述临床医生可用性、工作流整合和患者自我报告的结局。
提出的方法
- 将 AI Consult 部署为后台运行的安全网,在关键决策点通过交通灯界面(绿色/黄色/红色)呈现输出。
- 与 EMR 的异步、事件驱动集成,当用户离开关键字段时触发模型评审。
- 使用带本地情境和少样本示例的提示工程,以生成颜色、推理和建议行动。
- 比较在 15 家诊所、39,849 次就诊中,由 AI Consult 管理的就诊与未使用的就诊,对临床文档进行独立医生评审。
- 收集临床医生调查和常规随访电话,以评估可用性与患者自我报道的结果。
实验结果
研究问题
- RQ1AI Consult CDS 是否在真实的初级保健就诊中减少诊断与治疗错误?
- RQ2临床对齐的实施与主动部署如何影响工具的采用和有效性?
- RQ3对患者自我报告的结局和临床医生对照护质量的感知有何影响?
- RQ4在现实世界诊所中安全有效地采用基于LLM的 CDS 需要哪些关键因素?
主要发现
- 使用 AI Consult 的临床医生诊断错误减少了16%(NNT 18.1),治疗错误减少了13%(NNT 13.9)。
- 历史询问错误减少32%(NNT 11.3),调查/调查性错误减少10%(NNT 27.8)。
- 用非相对绝对值而言,AI Consult 将在 Penda 每年避免 22,000 次诊断错误和 29,000 次治疗错误。
- AI 组临床医生报告护理质量提高,75%表示影响显著,且所有受访者表示质量提高;没有患者自我报告的结局显示统计学显著差异。
- 基于 GPT-4.1 的评估显示的错误减少幅度大于医生评估者(例如治疗错误减少 22%、诊断错误减少 19%)。
- 未发现 AI Consult 的建议主动造成伤害的案例。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。