Skip to main content
QUICK REVIEW

[Paper Review] Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties

Jasmine Chiat Ling Ong, Liyuan Jin|arXiv (Cornell University)|Jan 29, 2024
Pharmacy and Medical Practices9 citations
TL;DR

This study introduces a Retrieval Augmented Generation (RAG) LLM framework as a clinical decision support system to improve detection of medication-related problems, comparing autonomous LLM use to a co-pilot setup with junior pharmacists across 12 specialties.

ABSTRACT

Importance: We introduce a novel Retrieval Augmented Generation (RAG)-Large Language Model (LLM) framework as a Clinical Decision Support Systems (CDSS) to support safe medication prescription. Objective: To evaluate the efficacy of LLM-based CDSS in correctly identifying medication errors in different patient case vignettes from diverse medical and surgical sub-disciplines, against a human expert panel derived ground truth. We compared performance for under 2 different CDSS practical healthcare integration modalities: LLM-based CDSS alone (fully autonomous mode) vs junior pharmacist + LLM-based CDSS (co-pilot, assistive mode). Design, Setting, and Participants: Utilizing a RAG model with state-of-the-art medically-related LLMs (GPT-4, Gemini Pro 1.0 and Med-PaLM 2), this study used 61 prescribing error scenarios embedded into 23 complex clinical vignettes across 12 different medical and surgical specialties. A multidisciplinary expert panel assessed these cases for Drug-Related Problems (DRPs) using the PCNE classification and graded severity / potential for harm using revised NCC MERP medication error index. We compared. Results RAG-LLM performed better compared to LLM alone. When employed in a co-pilot mode, accuracy, recall, and F1 scores were optimized, indicating effectiveness in identifying moderate to severe DRPs. The accuracy of DRP detection with RAG-LLM improved in several categories but at the expense of lower precision. Conclusions This study established that a RAG-LLM based CDSS significantly boosts the accuracy of medication error identification when used alongside junior pharmacists (co-pilot), with notable improvements in detecting severe DRPs. This study also illuminates the comparative performance of current state-of-the-art LLMs in RAG-based CDSS systems.

Motivation & Objective

  • Motivate the use of LLM-based CDSS to enhance safe medication prescription.
  • Evaluate whether a RAG-LLM CDSS improves identification of Drug-Related Problems (DRPs) against expert ground truth.
  • Assess performance differences between autonomous LLM use and a co-pilot model with junior pharmacists.
  • Investigate how state-of-the-art LLMs perform in a RAG-based CDSS across diverse medical and surgical specialties.

Proposed method

  • Utilize a Retrieval Augmented Generation (RAG) framework with GPT-4, Gemini Pro 1.0, and Med-PaLM 2 for medication safety decision support.
  • Embed 61 prescribing error scenarios into 23 complex vignettes spanning 12 specialties.
  • Have a multidisciplinary expert panel assess cases for Drug-Related Problems (DRPs) using PCNE classification and NCC MERP indices.
  • Compare RAG-LLM performance to LLM-alone (fully autonomous mode) and to a co-pilot mode with junior pharmacists.
  • Evaluate accuracy, recall, and F1 scores for DRP detection, noting trade-offs between improved accuracy and precision.

Experimental results

Research questions

  • RQ1Can a RAG-LLM CDSS accurately identify medication-related problems across diverse clinical specialties?
  • RQ2Does a co-pilot mode (junior pharmacist plus LLM) outperform autonomous LLM usage in detecting DRPs?
  • RQ3Which state-of-the-art LLMs (GPT-4, Gemini Pro, Med-PaLM 2) are most effective within a RAG-based CDSS for medication safety?
  • RQ4How do DRP severity and potential harm classifications influence CDSS performance across categories?

Key findings

  • RAG-LLM outperformed LLM-alone in detecting DRPs.
  • Co-pilot mode with junior pharmacists plus LLM achieved optimized accuracy, recall, and F1 scores for moderate to severe DRPs.
  • DRP detection accuracy improved in several categories with RAG-LLM, but precision decreased in some cases.
  • The study demonstrates the comparative performance of current state-of-the-art LLMs within RAG-based CDSS systems.
  • The RAG-LLM CDSS with co-pilot use notably enhances detection of severe DRPs across 12 specialties.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.