[论文解读] A Survey for Large Language Models in Biomedicine
本综述综合分析了484项生物医学大语言模型(LLM)研究,全面探讨了其在生物医学领域的应用、适应策略及挑战。研究突出展示了零样本性能在诊断和药物发现中的表现,微调在临床准确性方面的优势,并指出联邦学习与可解释人工智能是未来在隐私保护、可解释性及实际部署方面的重要方向。
Recent breakthroughs in large language models (LLMs) offer unprecedented natural language understanding and generation capabilities. However, existing surveys on LLMs in biomedicine often focus on specific applications or model architectures, lacking a comprehensive analysis that integrates the latest advancements across various biomedical domains. This review, based on an analysis of 484 publications sourced from databases including PubMed, Web of Science, and arXiv, provides an in-depth examination of the current landscape, applications, challenges, and prospects of LLMs in biomedicine, distinguishing itself by focusing on the practical implications of these models in real-world biomedical contexts. Firstly, we explore the capabilities of LLMs in zero-shot learning across a broad spectrum of biomedical tasks, including diagnostic assistance, drug discovery, and personalized medicine, among others, with insights drawn from 137 key studies. Then, we discuss adaptation strategies of LLMs, including fine-tuning methods for both uni-modal and multi-modal LLMs to enhance their performance in specialized biomedical contexts where zero-shot fails to achieve, such as medical question answering and efficient processing of biomedical literature. Finally, we discuss the challenges that LLMs face in the biomedicine domain including data privacy concerns, limited model interpretability, issues with dataset quality, and ethics due to the sensitive nature of biomedical data, the need for highly reliable model outputs, and the ethical implications of deploying AI in healthcare. To address these challenges, we also identify future research directions of LLM in biomedicine including federated learning methods to preserve data privacy and integrating explainable AI methodologies to enhance the transparency of LLMs.
研究动机与目标
- 提供对大型语言模型(LLMs)在多样化生物医学应用中的整体分析,超越狭窄聚焦领域。
- 识别并评估适用于单模态与多模态LLM的适应策略,如在临床与科研特定场景中的微调。
- 解决生物医学LLM中的关键挑战,包括数据隐私、模型可解释性、数据集质量及伦理化部署。
- 提出未来研究方向,重点关注联邦学习、可解释人工智能、高效微调及多模态模型融合。
- 通过严格验证与跨学科协作,指导LLM在临床工作流程中的负责任整合。
提出的方法
- 系统分析PubMed、Web of Science与arXiv收录的484篇文献,以描绘当前生物医学LLM的研究格局。
- 将LLM应用分类为诊断辅助、药物发现、个性化医学及生物医学文献处理。
- 评估零样本与微调后LLM在医学问答与基因组学分析等任务中的表现。
- 调查参数高效微调与多模态融合等适应技术,以整合文本、图像与结构化数据。
- 识别隐私保护方法,如联邦学习与差分隐私,以应对医疗数据敏感性问题。
- 探索注意力可视化、概念归因与LIME等可解释性技术,以提升模型透明度。

实验结果
研究问题
- RQ1通用大语言模型在零样本设置下,于诊断与药物发现等多样化生物医学任务中的表现如何?
- RQ2在临床决策支持等专业生物医学领域,提升LLM性能的最有效微调策略是什么?
- RQ3阻碍LLM在生物医学领域实际部署的主要挑战是什么,特别是数据隐私与模型可解释性方面?
- RQ4新兴技术如联邦学习与可解释人工智能如何缓解临床LLM应用中的伦理与技术障碍?
- RQ5为确保生物医学LLM的可靠性、公平性与全球适应性,哪些未来研究方向至关重要?
主要发现
- MedPaLM在医学问答任务中与临床专家达成92.9%的一致性,展示了其在复杂诊断推理中的强大零样本性能。
- 领域特定LLM如HuatuoGPT、ChatDoctor与BenTsao在医学对话与临床沟通任务中表现出高度可靠性。
- 微调显著提升了LLM在生物医学文献处理与医学问答等专业任务中的表现,而零样本性能在此类任务中不足。
- 整合文本、图像与结构化数据的多模态LLM在复杂生物医学分析中展现出更强能力,反映出模型架构的日益发展趋势。
- 联邦学习与差分隐私是保护数据隐私同时维持医疗环境中模型实用性的有前景方法。
- 可解释人工智能技术如注意力可视化与LIME可增强模型透明度,支持信任建立与临床采纳。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。