[Paper Review] Conversational Medical AI: Ready for Practice
This study evaluates Mo, a physician-supervised LLM-based conversational AI agent integrated into a real-world medical advice chat service at Alan. In a randomized controlled trial with 926 cases, patients reported significantly higher clarity (3.73/4 vs. 3.62, p < 0.05) and satisfaction (4.58/5 vs. 4.42, p < 0.05) with AI-assisted care, while safety was maintained with 95% of conversations rated 'good' or 'excellent' by physicians.
The shortage of doctors is creating a critical squeeze in access to medical expertise. While conversational Artificial Intelligence (AI) holds promise in addressing this problem, its safe deployment in patient-facing roles remains largely unexplored in real-world medical settings. We present the first large-scale evaluation of a physician-supervised LLM-based conversational agent in a real-world medical setting. Our agent, Mo, was integrated into an existing medical advice chat service. Over a three-week period, we conducted a randomized controlled experiment with 926 cases to evaluate patient experience and satisfaction. Among these, Mo handled 298 complete patient interactions, for which we report physician-assessed measures of safety and medical accuracy. Patients reported higher clarity of information (3.73 vs 3.62 out of 4, p < 0.05) and overall satisfaction (4.58 vs 4.42 out of 5, p < 0.05) with AI-assisted conversations compared to standard care, while showing equivalent levels of trust and perceived empathy. The high opt-in rate (81% among respondents) exceeded previous benchmarks for AI acceptance in healthcare. Physician oversight ensured safety, with 95% of conversations rated as "good" or "excellent" by general practitioners experienced in operating a medical advice chat service. Our findings demonstrate that carefully implemented AI medical assistants can enhance patient experience while maintaining safety standards through physician supervision. This work provides empirical evidence for the feasibility of AI deployment in healthcare communication and insights into the requirements for successful integration into existing healthcare services.
Motivation & Objective
- Address the global shortage of primary care physicians, especially in underserved rural areas, by deploying conversational AI to improve access to medical expertise.
- Evaluate the safety, accuracy, and patient experience of a real-world, physician-supervised LLM-based conversational AI agent in a live medical advice service.
- Assess patient satisfaction, trust, perceived empathy, and engagement in AI-assisted versus human-only medical consultations.
- Establish a framework for ethical, safe, and effective integration of conversational AI into existing healthcare workflows with clinical oversight.
- Provide empirical evidence for the feasibility of AI in patient-facing medical communication, informing future healthcare delivery models.
Proposed method
- Deployed Mo, a fine-tuned LLM-based conversational agent, within Alan’s existing physician-staffed medical chat service in France.
- Conducted a randomized controlled trial over three weeks, assigning 926 patients to either AI-assisted (Mo) or standard human-only consultations.
- Collected patient-reported outcomes via post-interaction surveys measuring clarity, satisfaction, trust, and perceived empathy.
- Implemented physician oversight: 298 complete AI-assisted conversations were clinically reviewed by experienced general practitioners for safety and accuracy.
- Used automated testing with simulated patient interactions to evaluate diagnostic reasoning and knowledge recall.
- Applied a comprehensive evaluation framework combining clinical assessment, real-world conversation analysis, and automated testing.
Experimental results
Research questions
- RQ1Does AI-assisted medical consultation improve patient-reported clarity of information and overall satisfaction compared to standard human-only care?
- RQ2How does patient trust and perceived empathy compare between AI-assisted and human-only medical consultations?
- RQ3What is the safety profile of a physician-supervised LLM-based conversational agent in real-world medical interactions?
- RQ4What is the level of patient engagement and opt-in rate for AI-assisted medical consultations in a real-world healthcare setting?
- RQ5How does the integration of conversational AI affect the quality and efficiency of medical advice delivery in primary care?
Key findings
- Patients reported significantly higher clarity of information in AI-assisted conversations (3.73 out of 4) compared to standard care (3.62 out of 4), with a p-value < 0.05.
- Overall patient satisfaction was higher in AI-assisted interactions (4.58 out of 5) than in standard care (4.42 out of 5), also with p < 0.05.
- Trust in the received information and perceived empathy were statistically equivalent between AI-assisted and human-only consultations.
- The opt-in rate for AI-assisted consultations reached 81% among respondents, exceeding previous benchmarks for AI acceptance in healthcare.
- Among 298 physician-reviewed conversations, 95% were rated as 'good' or 'excellent' for safety and medical accuracy, with no conversations deemed potentially dangerous.
- Patients engaged more actively with Mo, evidenced by shorter response times, indicating improved conversational flow and patient involvement.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.