Skip to main content
QUICK REVIEW

[논문 리뷰] Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties

Jasmine Chiat Ling Ong, Liyuan Jin|arXiv (Cornell University)|2024. 01. 29.
Pharmacy and Medical Practices인용 수 9
한 줄 요약

본 연구는 약물 관련 문제 탐지를 개선하기 위해 임상 의사결정 지원 시스템으로서 Retrieval Augmented Generation (RAG) LLM 프레임워크를 도입하고, 12개 전문 분야에 걸쳐 독립적 LLM 사용과 주니어 약사의 코파일럿 설정을 비교한다.

ABSTRACT

Importance: We introduce a novel Retrieval Augmented Generation (RAG)-Large Language Model (LLM) framework as a Clinical Decision Support Systems (CDSS) to support safe medication prescription. Objective: To evaluate the efficacy of LLM-based CDSS in correctly identifying medication errors in different patient case vignettes from diverse medical and surgical sub-disciplines, against a human expert panel derived ground truth. We compared performance for under 2 different CDSS practical healthcare integration modalities: LLM-based CDSS alone (fully autonomous mode) vs junior pharmacist + LLM-based CDSS (co-pilot, assistive mode). Design, Setting, and Participants: Utilizing a RAG model with state-of-the-art medically-related LLMs (GPT-4, Gemini Pro 1.0 and Med-PaLM 2), this study used 61 prescribing error scenarios embedded into 23 complex clinical vignettes across 12 different medical and surgical specialties. A multidisciplinary expert panel assessed these cases for Drug-Related Problems (DRPs) using the PCNE classification and graded severity / potential for harm using revised NCC MERP medication error index. We compared. Results RAG-LLM performed better compared to LLM alone. When employed in a co-pilot mode, accuracy, recall, and F1 scores were optimized, indicating effectiveness in identifying moderate to severe DRPs. The accuracy of DRP detection with RAG-LLM improved in several categories but at the expense of lower precision. Conclusions This study established that a RAG-LLM based CDSS significantly boosts the accuracy of medication error identification when used alongside junior pharmacists (co-pilot), with notable improvements in detecting severe DRPs. This study also illuminates the comparative performance of current state-of-the-art LLMs in RAG-based CDSS systems.

연구 동기 및 목표

  • 안전한 약물 처방을 강화하기 위해 LLM 기반 CDSS 사용의 필요성을 제기한다.
  • 전문가의 실제 기준(Ground truth)과 비교해 RAG-LLM CDSS가 DRP(약물 관련 문제) 식별을 향상시키는지 평가한다.
  • 자율 LLM 사용과 주니어 약사와의 코파일럿 모델 간의 성능 차이를 평가한다.
  • 다양한 의학 및 외과 전문 분야에서 최첨단 LLM이 RAG 기반 CDSS에서 어떻게 작동하는지 조사한다.

제안 방법

  • 의약 안전 의사결정 지원을 위해 GPT-4, Gemini Pro 1.0, 및 Med-PaLM 2를 활용한 Retrieval Augmented Generation (RAG) 프레임워크를 활용한다.
  • 12개 전문 분야에 걸친 23개의 복합 시나리오에 61건의 처방 오류 사례를 삽입한다.
  • 다분야 전문위원단이 PCNE 분류 및 NCC MERP 지수를 사용하여 DRP를 평가한다.
  • RAG-LLM의 성능을 LLM-단독(완전 자율 모드) 및 주니어 약사와의 코파일럿 모드와 비교한다.
  • DRP 탐지의 정확도, 재현율, F1 점수를 평가하고 향상된 정확도와 정밀도 간의 trade-off를 주의한다.

실험 결과

연구 질문

  • RQ1다양한 임상 전문 분야에서 RAG-LLM CDSS가 약물 관련 문제를 정확하게 식별할 수 있는가?
  • RQ2코파일럿 모드(주니어 약사와 LLM)가 DRP 탐지에서 자율 LLM 사용보다 우수한가?
  • RQ3RAG 기반 CDSS에서 약물 안전을 위해 가장 효과적인 최첨단 LLM(GPT-4, Gemini Pro, Med-PaLM 2)은 어느 것인가?
  • RQ4DRP 심각도와 잠재적 손상 분류가 카테고리 간 CDSS 성능에 어떤 영향을 미치는가?

주요 결과

  • RAG-LLM은 DRP 탐지에서 LLM-단독보다 우수했다.
  • 주니어 약사와 LLM의 코파일럿 모드는 중등도에서 중증 DRP에 대해 최적의 정확도, 재현율, F1 점수를 달성했다.
  • DRP 탐지 정확도는 RAG-LLM에서 여러 카테고리에서 향상되었으나 일부 경우에는 정밀도가 감소했다.
  • 이 연구는 RAG 기반 CDSS 시스템 내에서 현재 최첨단 LLM의 비교 성능을 보여준다.
  • 코파일럿 사용이 12개 전문 분야에 걸친 심각한 DRP 탐지 향상을 현저히 높인다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.