[Paper Review] Pre-trained Language Models in Biomedical Domain: A Systematic Survey
This survey comprehensively reviews biomedical pre-trained language models (PLMs), proposing a taxonomy, summarizing data sources, architectures, tasks, and benchmarks, and highlighting limitations and future trends. It includes textual, vision-language, and protein/DNA PLMs in biomedicine.
Pre-trained language models (PLMs) have been the de facto paradigm for most natural language processing (NLP) tasks. This also benefits biomedical domain: researchers from informatics, medicine, and computer science (CS) communities propose various PLMs trained on biomedical datasets, e.g., biomedical text, electronic health records, protein, and DNA sequences for various biomedical tasks. However, the cross-discipline characteristics of biomedical PLMs hinder their spreading among communities; some existing works are isolated from each other without comprehensive comparison and discussions. It expects a survey that not only systematically reviews recent advances of biomedical PLMs and their applications but also standardizes terminology and benchmarks. In this paper, we summarize the recent progress of pre-trained language models in the biomedical domain and their applications in biomedical downstream tasks. Particularly, we discuss the motivations and propose a taxonomy of existing biomedical PLMs. Their applications in biomedical downstream tasks are exhaustively discussed. At last, we illustrate various limitations and future trends, which we hope can provide inspiration for the future research of the research community.
Motivation & Objective
- Motivate the use of PLMs in the biomedical domain due to limited annotated data and abundant unlabeled data.
- Summarize the landscape of biomedical PLMs across data sources, architectures, and training paradigms.
- Propose a taxonomy of biomedical PLMs and provide practical resources and configurations for researchers.
Proposed method
- Survey the literature from Google Scholar, MedPub, Web of Science, and major conferences (ACL, EMNLP, NAACL, etc.).
- Categorize biomedical PLMs by data sources, model architecture, and pre-training/fine-tuning paradigms.
- Summarize downstream biomedical tasks (information extraction, classification, QA, dialogue, etc.) and corresponding PLM methods.
- Discuss limitations and future trends to guide future research in biomedical PLMs.

Experimental results
Research questions
- RQ1What data sources (texts, images, proteins/DNA, multi-modality) are used to pre-train biomedical PLMs?
- RQ2How are biomedical PLMs tailored from general-domain PLMs via domain adaptation and task adaptation?
- RQ3What are the main downstream tasks and benchmarks for biomedical PLMs, and how are they tackled?
- RQ4What are the current limitations of biomedical PLMs and what future directions are suggested?
- RQ5How do generative and vision-language or protein/DNA PLMs fit into biomedical research ecosystems?
Key findings
- Biomedical PLMs extend beyond text to vision-language models and sequence data (proteins/DNA).
- There is a proposed taxonomy categorizing PLMs by data sources, model architecture, and pre-training strategies.
- The survey consolidates resources, configurations, datasets, competitions, and venues to facilitate adoption by beginners.
- Biomedical PLMs face limitations and challenges that motivate future research directions and domain adaptation.
- This is among the first surveys to discuss vision-language and protein/DNA PLMs in biomedicine and to provide a broad, multi-scale overview.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.