Skip to main content
QUICK REVIEW

[Paper Review] AI for Biomedicine in the Era of Large Language Models

Zhenyu Bi, Sajib Acharjee Dip|arXiv (Cornell University)|Mar 23, 2024
Artificial Intelligence in Healthcare and Education5 citations
TL;DR

The paper surveys how large language models can be applied across three biomedical data types—text, biological sequences, and brain signals—and discusses challenges like trust, personalization, and multi-modal integration.

ABSTRACT

The capabilities of AI for biomedicine span a wide spectrum, from the atomic level, where it solves partial differential equations for quantum systems, to the molecular level, predicting chemical or protein structures, and further extending to societal predictions like infectious disease outbreaks. Recent advancements in large language models, exemplified by models like ChatGPT, have showcased significant prowess in natural language tasks, such as translating languages, constructing chatbots, and answering questions. When we consider biomedical data, we observe a resemblance to natural language in terms of sequences: biomedical literature and health records presented as text, biological sequences or sequencing data arranged in sequences, or sensor data like brain signals as time series. The question arises: Can we harness the potential of recent large language models to drive biomedical knowledge discoveries? In this survey, we will explore the application of large language models to three crucial categories of biomedical data: 1) textual data, 2) biological sequences, and 3) brain signals. Furthermore, we will delve into large language model challenges in biomedical research, including ensuring trustworthiness, achieving personalization, and adapting to multi-modal data representation

Motivation & Objective

  • Motivate the exploration of large language models for biomedical knowledge discovery across textual data, biological sequences, and brain signals.
  • Survey existing biomedical LLMs and architectures tailored to different data modalities.
  • Identify challenges and considerations for trustworthy, personalized, and multi-modal biomedical AI.
  • Highlight applications and downstream tasks enabled by biomedical LLMs in clinical and research settings.

Proposed method

  • Review and synthesize literature on biomedical LLMs across three data categories: textual data, biological sequences, and brain signals.
  • Summarize representative models and pre-training strategies in each modality (e.g., SciBERT, BioBERT, BioGPT, GenSLMs, DNABERT, RNABERT, ESM, ProtST, etc.).
  • Discuss applications in information extraction, question answering, relation extraction, and protein/RNA sequence understanding and generation.
  • Outline challenges including trustworthiness, personalization, and multi-modal data representation.
Figure 1. Overview of applications, models, and downstream tasks for biomedical LLMs on textual data.
Figure 1. Overview of applications, models, and downstream tasks for biomedical LLMs on textual data.

Experimental results

Research questions

  • RQ1What are the current capabilities and limitations of LLMs when applied to biomedical textual data, sequences, and brain signals?
  • RQ2Which models and pre-training approaches are most effective for each biomedical data modality?
  • RQ3What challenges must be addressed to ensure trustworthy, personalized, and multi-modal LLM-driven biomedical research and practice?

Key findings

  • Biomedical LLMs span textual data, sequences (DNA, RNA, proteins, multi-omics), and brain signals with specialized models and pre-training regimes.
  • A wide range of domain-specific models (e.g., SciBERT, BioBERT, PubMedBERT, BioLinkBERT, Galactica, BioGPT, DoT5, GenSLMs, DNABERT, RNABERT, RNA-MSM, ESM/ESM-2, ProtST) have achieved state-of-the-art or near state-of-the-art performance on diverse tasks across the three data categories.
  • Applications include information extraction, question answering, relation extraction, and advanced sequence/function prediction and design for biomedical discovery.
  • The paper emphasizes challenges such as ensuring trustworthiness, enabling personalization, and adapting models to multi-modal biomedical data.
Figure 2. Overview of models and applications for Genomic LLMs on biological sequences.
Figure 2. Overview of models and applications for Genomic LLMs on biological sequences.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.