Skip to main content
QUICK REVIEW

[Paper Review] A Survey for Large Language Models in Biomedicine

Chong Wang, Mengyao Li|arXiv (Cornell University)|Aug 29, 2024
Machine Learning in Healthcare4 citations
TL;DR

This survey synthesizes 484 biomedical LLM studies to comprehensively analyze applications, adaptation strategies, and challenges in biomedicine. It highlights zero-shot performance in diagnostics and drug discovery, fine-tuning for clinical accuracy, and identifies federated learning and explainable AI as key future directions for privacy, interpretability, and real-world deployment.

ABSTRACT

Recent breakthroughs in large language models (LLMs) offer unprecedented natural language understanding and generation capabilities. However, existing surveys on LLMs in biomedicine often focus on specific applications or model architectures, lacking a comprehensive analysis that integrates the latest advancements across various biomedical domains. This review, based on an analysis of 484 publications sourced from databases including PubMed, Web of Science, and arXiv, provides an in-depth examination of the current landscape, applications, challenges, and prospects of LLMs in biomedicine, distinguishing itself by focusing on the practical implications of these models in real-world biomedical contexts. Firstly, we explore the capabilities of LLMs in zero-shot learning across a broad spectrum of biomedical tasks, including diagnostic assistance, drug discovery, and personalized medicine, among others, with insights drawn from 137 key studies. Then, we discuss adaptation strategies of LLMs, including fine-tuning methods for both uni-modal and multi-modal LLMs to enhance their performance in specialized biomedical contexts where zero-shot fails to achieve, such as medical question answering and efficient processing of biomedical literature. Finally, we discuss the challenges that LLMs face in the biomedicine domain including data privacy concerns, limited model interpretability, issues with dataset quality, and ethics due to the sensitive nature of biomedical data, the need for highly reliable model outputs, and the ethical implications of deploying AI in healthcare. To address these challenges, we also identify future research directions of LLM in biomedicine including federated learning methods to preserve data privacy and integrating explainable AI methodologies to enhance the transparency of LLMs.

Motivation & Objective

  • To provide a holistic analysis of large language models (LLMs) across diverse biomedical applications beyond narrow focus areas.
  • To identify and evaluate adaptation strategies such as fine-tuning for unimodal and multimodal LLMs in specialized clinical and research contexts.
  • To address critical challenges in biomedical LLMs, including data privacy, model interpretability, dataset quality, and ethical deployment.
  • To propose future research directions focused on federated learning, explainable AI, efficient fine-tuning, and multimodal model fusion.
  • To guide responsible integration of LLMs into clinical workflows through rigorous validation and interdisciplinary collaboration.

Proposed method

  • Systematic analysis of 484 publications from PubMed, Web of Science, and arXiv to map the current landscape of biomedical LLMs.
  • Categorization of LLM applications into diagnostic assistance, drug discovery, personalized medicine, and biomedical literature processing.
  • Evaluation of zero-shot and fine-tuned LLM performance across tasks such as medical question answering and genomics analysis.
  • Survey of adaptation techniques including parameter-efficient fine-tuning and multimodal fusion for integrating text, images, and structured data.
  • Identification of privacy-preserving methods like federated learning and differential privacy to address data sensitivity in healthcare.
  • Exploration of interpretability techniques such as attention visualization, concept attribution, and LIME to improve model transparency.
Figure 1: Chronological overview of LLMs and their variants in biomedical applications from 2019 to 2024. The timeline illustrates the evolution of both unimodal (top) and multimodal (bottom) models, highlighting key developments across different model architectures including LLAMA, GPT, BERT, BaiCh
Figure 1: Chronological overview of LLMs and their variants in biomedical applications from 2019 to 2024. The timeline illustrates the evolution of both unimodal (top) and multimodal (bottom) models, highlighting key developments across different model architectures including LLAMA, GPT, BERT, BaiCh

Experimental results

Research questions

  • RQ1How do general-purpose LLMs perform in zero-shot settings across diverse biomedical tasks such as diagnosis and drug discovery?
  • RQ2What are the most effective fine-tuning strategies for enhancing LLM performance in specialized biomedical domains like clinical decision support?
  • RQ3What are the primary challenges hindering the real-world deployment of LLMs in biomedicine, particularly concerning data privacy and model interpretability?
  • RQ4How can emerging techniques like federated learning and explainable AI mitigate ethical and technical barriers in clinical LLM applications?
  • RQ5What future research directions are essential for ensuring the reliability, fairness, and global adaptability of biomedical LLMs?

Key findings

  • MedPaLM achieved 92.9% agreement with clinical experts in medical question answering, demonstrating strong zero-shot performance in complex diagnostic reasoning.
  • Domain-specific LLMs such as HuatuoGPT, ChatDoctor, and BenTsao show high reliability in medical dialogue and clinical communication tasks.
  • Fine-tuning significantly improves LLM performance in specialized tasks like biomedical literature processing and medical QA where zero-shot performance is insufficient.
  • Multimodal LLMs integrating text, images, and structured data show enhanced capability in complex biomedical analysis, reflecting a growing trend in model architecture.
  • Federated learning and differential privacy are promising approaches to preserving data privacy while maintaining model utility in healthcare settings.
  • Explainable AI techniques like attention visualization and LIME can enhance model transparency, supporting trust and clinical adoption.
Figure 2: Trends and distribution of LLM research papers in biomedical fields from 2018 to 2024. (a) Temporal analysis of LLM research papers, showing quarterly publication counts. A surge in publications is evident beginning in 2021, reflecting growing interest and investment in applying LLMs to bi
Figure 2: Trends and distribution of LLM research papers in biomedical fields from 2018 to 2024. (a) Temporal analysis of LLM research papers, showing quarterly publication counts. A surge in publications is evident beginning in 2021, reflecting growing interest and investment in applying LLMs to bi

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.