Skip to main content
QUICK REVIEW

[论文解读] Pre-trained Language Models in Biomedical Domain: A Systematic Survey

Benyou Wang, Qianqian Xie|arXiv (Cornell University)|Oct 11, 2021
Topic Modeling参考文献 295被引用 34
一句话总结

本综述全面评估生物医学预训练语言模型(PLMs),提出一种分类法,概括数据源、架构、任务和基准,并突出其局限性和未来趋势。它涵盖生物医学中的文本、视觉-语言以及蛋白质/DNA PLMs。

ABSTRACT

Pre-trained language models (PLMs) have been the de facto paradigm for most natural language processing (NLP) tasks. This also benefits biomedical domain: researchers from informatics, medicine, and computer science (CS) communities propose various PLMs trained on biomedical datasets, e.g., biomedical text, electronic health records, protein, and DNA sequences for various biomedical tasks. However, the cross-discipline characteristics of biomedical PLMs hinder their spreading among communities; some existing works are isolated from each other without comprehensive comparison and discussions. It expects a survey that not only systematically reviews recent advances of biomedical PLMs and their applications but also standardizes terminology and benchmarks. In this paper, we summarize the recent progress of pre-trained language models in the biomedical domain and their applications in biomedical downstream tasks. Particularly, we discuss the motivations and propose a taxonomy of existing biomedical PLMs. Their applications in biomedical downstream tasks are exhaustively discussed. At last, we illustrate various limitations and future trends, which we hope can provide inspiration for the future research of the research community.

研究动机与目标

  • 由于标注数据有限、未标注数据丰富,推动在生物医学领域使用PLMs。
  • 总结跨数据源、架构和训练范式的生物医学PLMs的全景。
  • 提出生物医学PLMs的分类法,并为研究者提供实用资源和配置。

提出的方法

  • 对 Google Scholar、MedPub、Web of Science 以及主要会议(ACL、EMNLP、NAACL 等)中的文献进行调研。
  • 按数据源、模型架构和预训练/微调范式对生物医学PLMs进行分类。
  • 总结下游生物医学任务(信息抽取、分类、问答、对话等)及相应的PLM方法。
  • 讨论局限性和未来趋势,以引导生物医学PLMs的未来研究。
Figure 1 . Overview of selected released Biomedical pre-trained language models. One can see a more detailed list in Sec. 3 . Note that there is a BERT-like language model embedded in the overall architecture of AlphaFold 2.
Figure 1 . Overview of selected released Biomedical pre-trained language models. One can see a more detailed list in Sec. 3 . Note that there is a BERT-like language model embedded in the overall architecture of AlphaFold 2.

实验结果

研究问题

  • RQ1用于预训练生物医学PLMs的数据源有哪些(文本、图像、蛋白质/DNA、多模态)?
  • RQ2通过领域自适应和任务自适应,生物医学PLMs如何从通用领域PLMs定制?
  • RQ3生物医学PLMs的主要下游任务和基准是什么,如何应对?
  • RQ4生物医学PLMs当前的局限性与未来方向有哪些?
  • RQ5生成式以及视觉-语言或蛋白质/DNA PLMs如何融入生物医学研究生态系统?

主要发现

  • 生物医学PLMs不仅限于文本,还扩展到视觉-语言模型和序列数据(蛋白质/DNA)。
  • 提出了一种按数据源、模型架构和预训练策略对PLMs进行分类的分类法。
  • 本综述整合了资源、配置、数据集、竞赛和场馆信息,便于初学者采用。
  • 生物医学PLMs面临的局限性和挑战促使未来的研究方向和领域自适应。
  • 这是少数几份讨论生物医学中的视觉-语言和蛋白质/DNA PLMs并提供广泛的多尺度综述的前沿调查之一。
Figure 2 . Parallel development of general and biomedical pre-trained language models. The time is determined by the released date of the paper, for example, in arXiv. General pre-trained language models are shown below in the timeline, and biomedical pre-trained language models are shown above the
Figure 2 . Parallel development of general and biomedical pre-trained language models. The time is determined by the released date of the paper, for example, in arXiv. General pre-trained language models are shown below in the timeline, and biomedical pre-trained language models are shown above the

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。