[Paper Review] Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties
This study introduces a Retrieval Augmented Generation (RAG) LLM framework as a clinical decision support system to improve detection of medication-related problems, comparing autonomous LLM use to a co-pilot setup with junior pharmacists across 12 specialties.
Importance: We introduce a novel Retrieval Augmented Generation (RAG)-Large Language Model (LLM) framework as a Clinical Decision Support Systems (CDSS) to support safe medication prescription. Objective: To evaluate the efficacy of LLM-based CDSS in correctly identifying medication errors in different patient case vignettes from diverse medical and surgical sub-disciplines, against a human expert panel derived ground truth. We compared performance for under 2 different CDSS practical healthcare integration modalities: LLM-based CDSS alone (fully autonomous mode) vs junior pharmacist + LLM-based CDSS (co-pilot, assistive mode). Design, Setting, and Participants: Utilizing a RAG model with state-of-the-art medically-related LLMs (GPT-4, Gemini Pro 1.0 and Med-PaLM 2), this study used 61 prescribing error scenarios embedded into 23 complex clinical vignettes across 12 different medical and surgical specialties. A multidisciplinary expert panel assessed these cases for Drug-Related Problems (DRPs) using the PCNE classification and graded severity / potential for harm using revised NCC MERP medication error index. We compared. Results RAG-LLM performed better compared to LLM alone. When employed in a co-pilot mode, accuracy, recall, and F1 scores were optimized, indicating effectiveness in identifying moderate to severe DRPs. The accuracy of DRP detection with RAG-LLM improved in several categories but at the expense of lower precision. Conclusions This study established that a RAG-LLM based CDSS significantly boosts the accuracy of medication error identification when used alongside junior pharmacists (co-pilot), with notable improvements in detecting severe DRPs. This study also illuminates the comparative performance of current state-of-the-art LLMs in RAG-based CDSS systems.
Motivation & Objective
- Motivate the use of LLM-based CDSS to enhance safe medication prescription.
- Evaluate whether a RAG-LLM CDSS improves identification of Drug-Related Problems (DRPs) against expert ground truth.
- Assess performance differences between autonomous LLM use and a co-pilot model with junior pharmacists.
- Investigate how state-of-the-art LLMs perform in a RAG-based CDSS across diverse medical and surgical specialties.
Proposed method
- Utilize a Retrieval Augmented Generation (RAG) framework with GPT-4, Gemini Pro 1.0, and Med-PaLM 2 for medication safety decision support.
- Embed 61 prescribing error scenarios into 23 complex vignettes spanning 12 specialties.
- Have a multidisciplinary expert panel assess cases for Drug-Related Problems (DRPs) using PCNE classification and NCC MERP indices.
- Compare RAG-LLM performance to LLM-alone (fully autonomous mode) and to a co-pilot mode with junior pharmacists.
- Evaluate accuracy, recall, and F1 scores for DRP detection, noting trade-offs between improved accuracy and precision.
Experimental results
Research questions
- RQ1Can a RAG-LLM CDSS accurately identify medication-related problems across diverse clinical specialties?
- RQ2Does a co-pilot mode (junior pharmacist plus LLM) outperform autonomous LLM usage in detecting DRPs?
- RQ3Which state-of-the-art LLMs (GPT-4, Gemini Pro, Med-PaLM 2) are most effective within a RAG-based CDSS for medication safety?
- RQ4How do DRP severity and potential harm classifications influence CDSS performance across categories?
Key findings
- RAG-LLM outperformed LLM-alone in detecting DRPs.
- Co-pilot mode with junior pharmacists plus LLM achieved optimized accuracy, recall, and F1 scores for moderate to severe DRPs.
- DRP detection accuracy improved in several categories with RAG-LLM, but precision decreased in some cases.
- The study demonstrates the comparative performance of current state-of-the-art LLMs within RAG-based CDSS systems.
- The RAG-LLM CDSS with co-pilot use notably enhances detection of severe DRPs across 12 specialties.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.