[Paper Review] AI-based Clinical Decision Support for Primary Care: A Real-World Study
This study evaluates an LLM-based clinical decision support tool (AI Consult) deployed in Nairobi, Kenya, showing reduced clinical errors in live care and positive clinician feedback, with no significant difference in patient-reported outcomes.
We evaluate the impact of large language model-based clinical decision support in live care. In partnership with Penda Health, a network of primary care clinics in Nairobi, Kenya, we studied AI Consult, a tool that serves as a safety net for clinicians by identifying potential documentation and clinical decision-making errors. AI Consult integrates into clinician workflows, activating only when needed and preserving clinician autonomy. We conducted a quality improvement study, comparing outcomes for 39,849 patient visits performed by clinicians with or without access to AI Consult across 15 clinics. Visits were rated by independent physicians to identify clinical errors. Clinicians with access to AI Consult made relatively fewer errors: 16% fewer diagnostic errors and 13% fewer treatment errors. In absolute terms, the introduction of AI Consult would avert diagnostic errors in 22,000 visits and treatment errors in 29,000 visits annually at Penda alone. In a survey of clinicians with AI Consult, all clinicians said that AI Consult improved the quality of care they delivered, with 75% saying the effect was "substantial". These results required a clinical workflow-aligned AI Consult implementation and active deployment to encourage clinician uptake. We hope this study demonstrates the potential for LLM-based clinical decision support tools to reduce errors in real-world settings and provides a practical framework for advancing responsible adoption.
Motivation & Objective
- Assess whether an LLM-based CDS tool reduces clinical documentation and decision-making errors in primary care.
- Evaluate how clinically-aligned implementation and active deployment influence uptake and effectiveness.
- Characterize clinician usability, workflow integration, and patient-reported outcomes in a real-world setting.
Proposed method
- Deploy AI Consult as a safety-net that runs in the background and surfaces outputs via a traffic-light interface (green/yellow/red) at key decision points.
- Use asynchronous, event-driven integration with the EMR to trigger model reviews when users navigate away from critical fields.
- Apply prompt engineering with local context and few-shot examples to produce color, rationale, and suggested actions.
- Compare visits managed with AI Consult to those without across 15 clinics and 39,849 visits, with independent physician reviews of clinical documentation.
- Collect clinician surveys and routine follow-up calls to assess usability and patient-reported outcomes.
Experimental results
Research questions
- RQ1Does the AI Consult CDS reduce diagnostic and treatment errors in live primary care visits?
- RQ2How does clinically-aligned implementation and active deployment affect tool uptake and effectiveness?
- RQ3What is the impact on patient-reported outcomes and clinician-perceived quality of care?
- RQ4What factors are critical for the safe and effective adoption of LLM-based CDS in real-world clinics?
Key findings
- Clinicians with AI Consult had 16% fewer diagnostic errors (NNT 18.1) and 13% fewer treatment errors (NNT 13.9).
- There were 32% reductions in history-taking errors (NNT 11.3) and 10% reductions in investigation errors (NNT 27.8).
- In absolute terms, AI Consult would avert diagnostic errors in 22,000 visits and treatment errors in 29,000 visits annually at Penda.
- Clinicians in the AI group reported improved quality of care, with 75% saying the effect was substantial, and all respondents indicating improved quality; no patient-reported outcomes showed a statistically significant difference.
- GPT-4.1-based evaluations suggested larger error reductions than physician evaluators (e.g., 22% treatment and 19% diagnostic error reductions).
- No cases were found where AI Consult advice actively caused harm.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.