Skip to main content
QUICK REVIEW

[论文解读] Pre-trained Language Models in Biomedical Domain: A Survey from Multiscale Perspective

Benyou Wang, Qianqian Xie|arXiv (Cornell University)|Oct 11, 2021
Artificial Intelligence in Healthcare and Education被引用 2
一句话总结

本综述对生物医学领域中的预训练语言模型(PLMs)进行了系统性、多尺度的回顾,对现有模型进行分类,统一术语和基准,并分析其在下游任务中的应用。该综述提出了一套统一的分类体系,指出了关键局限性及未来跨学科发展的研究方向。

ABSTRACT

Pre-trained language models have been the de facto paradigm for most natural language processing (NLP) tasks. In the biomedical domain, which also benefits from NLP techniques, various pre-trained language models were proposed by leveraging domain datasets including biomedical literature, biomedical social medial, electronic health records, and other biological sequences. Large amounts of efforts have been explored on applying these biomedical pre-trained language models to downstream biomedical tasks, from informatics, medicine, and computer science (CS) communities. However, it seems that the vast majority of existing works are isolated from each other probably because of the cross-discipline characteristics. It is expected to propose a survey that not only systematically reviews recent advances of biomedical pre-trained language models and their applications but also standardizes terminology, taxonomy, and benchmarks. Therefore, this paper summarizes the recent progress of pre-trained language models used in the biomedical domain. Particularly, an overview and taxonomy of existing biomedical pre-trained language models as well as their applications in biomedical downstream tasks are exhaustively discussed. At last, we illustrate various limitations and future trends, which we hope can provide inspiration for the future research.

研究动机与目标

  • 系统性回顾多尺度下生物医学预训练语言模型的最新进展。
  • 统一生物医学PLM的术语、分类体系和评估基准。
  • 弥合生物医学、信息学与计算机科学领域中孤立的研究工作之间的差距。
  • 识别生物医学自然语言处理中现有局限性及新兴研究趋势,以指导未来研究。

提出的方法

  • 基于训练数据来源(如生物医学文献、电子健康记录EHRs、生物序列)对生物医学PLM进行分类。
  • 构建统一的分类体系,按模型架构、预训练目标和领域特异性对模型进行分类。
  • 分析模型在下游生物医学任务中的应用,包括命名实体识别、关系抽取和问答系统。
  • 整合现有基准测试与评估协议,以支持可复现性和可比性。
  • 识别模型泛化与领域适应方面存在的重复性方法论缺陷与挑战。

实验结果

研究问题

  • RQ1生物医学预训练语言模型在架构、预训练数据和目标方面如何分类?
  • RQ2现有生物医学PLM在设计与应用方面的主要差异与相似之处是什么?
  • RQ3生物医学PLM在疾病预测和药物反应建模等多样化下游任务中的表现如何?
  • RQ4生物医学NLP研究中在术语、基准测试与评估方面存在哪些标准化挑战?
  • RQ5生物医学PLM开发中的主要局限性与新兴研究趋势是什么?

主要发现

  • 已建立涵盖数据来源、架构与预训练目标的生物医学预训练语言模型综合分类体系。
  • 综述指出,生物医学NLP研究在术语、基准与评估协议方面缺乏标准化。
  • 现有模型在命名实体识别与关系抽取等任务中表现优异,尤其在微调领域特定数据后。
  • 尽管已有进展,但生物医学、信息学与计算机科学领域间的研究孤立状态仍阻碍跨学科发展。
  • 未来研究应聚焦于提升模型泛化能力、减少数据偏差,并建立共享基准以实现可复现性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。