Skip to main content
QUICK REVIEW

[論文レビュー] Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties

Jasmine Chiat Ling Ong, Liyuan Jin|arXiv (Cornell University)|Jan 29, 2024
Pharmacy and Medical Practices被引用数 9
ひとこと要約

この研究は、薬物関連の問題の検出を改善する臨床意思決定支援システムとして Retrieval Augmented Generation (RAG) LLM framework を導入し、自律型 LLM の使用と junior pharmacists を用いた co-pilot セットアップを 12 の専門分野に跨って比較します。

ABSTRACT

Importance: We introduce a novel Retrieval Augmented Generation (RAG)-Large Language Model (LLM) framework as a Clinical Decision Support Systems (CDSS) to support safe medication prescription. Objective: To evaluate the efficacy of LLM-based CDSS in correctly identifying medication errors in different patient case vignettes from diverse medical and surgical sub-disciplines, against a human expert panel derived ground truth. We compared performance for under 2 different CDSS practical healthcare integration modalities: LLM-based CDSS alone (fully autonomous mode) vs junior pharmacist + LLM-based CDSS (co-pilot, assistive mode). Design, Setting, and Participants: Utilizing a RAG model with state-of-the-art medically-related LLMs (GPT-4, Gemini Pro 1.0 and Med-PaLM 2), this study used 61 prescribing error scenarios embedded into 23 complex clinical vignettes across 12 different medical and surgical specialties. A multidisciplinary expert panel assessed these cases for Drug-Related Problems (DRPs) using the PCNE classification and graded severity / potential for harm using revised NCC MERP medication error index. We compared. Results RAG-LLM performed better compared to LLM alone. When employed in a co-pilot mode, accuracy, recall, and F1 scores were optimized, indicating effectiveness in identifying moderate to severe DRPs. The accuracy of DRP detection with RAG-LLM improved in several categories but at the expense of lower precision. Conclusions This study established that a RAG-LLM based CDSS significantly boosts the accuracy of medication error identification when used alongside junior pharmacists (co-pilot), with notable improvements in detecting severe DRPs. This study also illuminates the comparative performance of current state-of-the-art LLMs in RAG-based CDSS systems.

研究の動機と目的

  • LLM ベースの CDSS を活用して安全な薬物処方を強化する動機づけ。
  • RAG-LLM CDSS が DRPs の特定を専門家のグラウンドトゥルースと比較して改善しているかを評価する。
  • 自律 LLM の使用と junior pharmacists を伴う co-pilot モデルとの性能差を評価する。
  • 先端的な LLM が多様な医療・外科系専門分野を横断する RAG ベースの CDSS でどのように機能するかを調査する。

提案手法

  • GPT-4、Gemini Pro 1.0、および Med-PaLM 2 を用いた RAG フレームワークを薬物安全性意思決定支援に適用する。
  • 12 の専門分野に跨る 61 の処方エラーシナリオを 23 の複雑なビネットに埋め込む。
  • 多職種の専門家パネルが PCNE 分类と NCC MERP 指標を用いて DRP の症例を評価する。
  • RAG-LLM の性能を LLM-のみ(完全自律モード)および junior pharmacists を伴う co-pilot モードと比較する。
  • DRP 検出の正確性・再現性(recall)・F1 スコアを評価し、精度が向上する一方で適合率が低下するトレードオフに留意する。

実験結果

リサーチクエスチョン

  • RQ1RAG-LLM CDSS は多様な臨床専門分野で薬物関連の問題を正確に識別できるか。
  • RQ2コ-pilot モード(junior pharmacist + LLM)は自律 LLM の使用を上回る DRP の検出性能を示すか。
  • RQ3RAG ベースの CDSS において、最先端の LLM(GPT-4、Gemini Pro、Med-PaLM 2)はどれが最も有効か。
  • RQ4DRP の Severity および潜在的な有害性分類は CDSS の性能をカテゴリ間でどう影響するか。

主な発見

  • RAG-LLM は DRP の検出において LLM-alone を上回った。
  • junior pharmacists を伴う co-pilot モードは中等度から重度の DRP について正確性・再現性・F1 スコアを最適化した。
  • RAG-LLM によっていくつかのカテゴリで DRP 検出の正確性が向上したが、いくつかのケースで適合率が低下した。
  • 本研究は現在の最先端 LLM の RAG ベース CDSS 系統における比較性能を示す。
  • co-pilot 使用による RAG-LLM CDSS は 12 の専門分野で重度 DRP の検出を著しく改善する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。