[论文解读] Stochastic Parrots or ICU Experts? Large Language Models in Critical Care Medicine: A Scoping Review
本综述性研究探讨了大型语言模型(LLMs)在重症监护医学(CCM)中的应用,分析了2019–2024年间24项研究。研究发现,LLMs在临床决策支持、病历记录和医学教育方面展现出潜力,但面临幻觉、可解释性差、偏见和伦理风险等挑战,亟需提升可靠性、知识整合能力以及建立更完善的伦理框架。
With the rapid development of artificial intelligence (AI), large language models (LLMs) have shown strong capabilities in natural language understanding, reasoning, and generation, attracting amounts of research interest in applying LLMs to health and medicine. Critical care medicine (CCM) provides diagnosis and treatment for critically ill patients who often require intensive monitoring and interventions in intensive care units (ICUs). Can LLMs be applied to CCM? Are LLMs just like stochastic parrots or ICU experts in assisting clinical decision-making? This scoping review aims to provide a panoramic portrait of the application of LLMs in CCM. Literature in seven databases, including PubMed, Embase, Scopus, Web of Science, CINAHL, IEEE Xplore, and ACM Digital Library, were searched from January 1, 2019, to June 10, 2024. Peer-reviewed journal and conference articles that discussed the application of LLMs in critical care settings were included. From an initial 619 articles, 24 were selected for final review. This review grouped applications of LLMs in CCM into three categories: clinical decision support, medical documentation and reporting, and medical education and doctor-patient communication. LLMs have advantages in handling unstructured data and do not require manual feature engineering. Meanwhile, applying LLMs to CCM faces challenges, including hallucinations, poor interpretability, bias and alignment challenges, and privacy and ethics issues. Future research should enhance model reliability and interpretability, integrate up-to-date medical knowledge, and strengthen privacy and ethical guidelines. As LLMs evolve, they could become key tools in CCM to help improve patient outcomes and optimize healthcare delivery. This study is the first review of LLMs in CCM, aiding researchers, clinicians, and policymakers to understand the current status and future potentials of LLMs in CCM.
研究动机与目标
- 绘制大型语言模型(LLMs)在重症监护医学(CCM)中应用的现状图景。
- 识别LLMs应用的关键领域,包括临床决策支持、病历记录和医学教育。
- 评估在高风险ICU环境中部署LLMs的益处与风险。
- 突出临床LLM应用中幻觉、可解释性差、偏见和伦理问题等关键挑战。
- 通过识别优先研究方向,为未来研究提供指导:模型可靠性、实时知识整合以及健全的隐私与伦理标准。
提出的方法
- 在七个数据库中系统性检索:PubMed、Embase、Scopus、Web of Science、CINAHL、IEEE Xplore和ACM数字图书馆。
- 纳入标准:2019年1月1日至2024年6月10日期间发表的同行评审期刊文章和会议论文,聚焦LLMs在CCM中的应用。
- 通过题名/摘要和全文评审筛选619篇文献,最终纳入24项研究进行分析。
- 将LLM应用分类为三个领域:临床决策支持、病历记录与报告、医学教育及医患沟通。
- 主题化综合研究发现,重点强调LLM在ICU中部署所面临的工程技术、伦理和临床挑战。
- 基于在可靠性、可解释性和伦理治理方面识别出的研究空白,提出未来研究建议。
实验结果
研究问题
- RQ1近期文献中报告的大型语言模型在重症监护医学中的主要应用是什么?
- RQ2大型语言模型在重症监护环境中处理非结构化临床数据的表现如何?
- RQ3在ICU环境中部署LLMs所关联的主要技术和伦理挑战有哪些?
- RQ4LLMs在多大程度上提升了临床决策、病历记录效率或医学教育的质量?
- RQ5为提升LLMs在CCM中的安全性、可靠性与伦理使用,未来研究应关注哪些方向?
主要发现
- LLMs在处理非结构化临床数据方面展现出强大能力,并减少了CCM应用中对手动特征工程的需求。
- CCM中LLM应用的大多数案例被归类为三个领域:临床决策支持(如诊断与治疗建议)、病历记录(如自动病历生成)以及医学教育(如培训与沟通工具)。
- 幻觉——即生成事实错误或虚构的临床信息——在多个研究中被列为显著风险,尤其在诊断和治疗建议任务中更为突出。
- 模型可解释性差以及推理过程缺乏透明度,被一致认为是临床信任与采纳的主要障碍。
- 训练数据中的偏见以及与临床指南的不一致,被识别为影响LLM输出公平性与安全性的关键风险。
- 隐私和伦理问题,包括患者数据泄露风险以及监管监督的缺失,频繁被提及为现实部署中的关键障碍。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。