Skip to main content
QUICK REVIEW

[논문 리뷰] Pre-trained Language Models in Biomedical Domain: A Systematic Survey

Benyou Wang, Qianqian Xie|arXiv (Cornell University)|2021. 10. 11.
Topic Modeling참고 문헌 295인용 수 34
한 줄 요약

본 고찰은 생물의학 분야의 사전 학습 언어 모델(PLMs)을 포괄적으로 검토하고, 분류 체계(계통학)를 제안하며, 데이터 소스, 아키텍처, 태스크, 벤치마크를 요약하고 한계와 향후 경향을 강조합니다. 또한 생물의학에서 텍스트 기반 PLM뿐 아니라 비전-언어 및 단백질/DNA PLMs를 포함합니다.

ABSTRACT

Pre-trained language models (PLMs) have been the de facto paradigm for most natural language processing (NLP) tasks. This also benefits biomedical domain: researchers from informatics, medicine, and computer science (CS) communities propose various PLMs trained on biomedical datasets, e.g., biomedical text, electronic health records, protein, and DNA sequences for various biomedical tasks. However, the cross-discipline characteristics of biomedical PLMs hinder their spreading among communities; some existing works are isolated from each other without comprehensive comparison and discussions. It expects a survey that not only systematically reviews recent advances of biomedical PLMs and their applications but also standardizes terminology and benchmarks. In this paper, we summarize the recent progress of pre-trained language models in the biomedical domain and their applications in biomedical downstream tasks. Particularly, we discuss the motivations and propose a taxonomy of existing biomedical PLMs. Their applications in biomedical downstream tasks are exhaustively discussed. At last, we illustrate various limitations and future trends, which we hope can provide inspiration for the future research of the research community.

연구 동기 및 목표

  • 주석 데이터가 제한되고 미라벨 데이터가 풍부한 생물의학 영역에서 PLMs 활용의 필요성을 제시한다.
  • 데이터 소스, 아키텍처, 학습 패러다임 전반에 걸친 생물의학 PLMs의 현황을 요약한다.
  • 생물의학 PLMs의 분류 체계를 제시하고 연구자들을 위한 실용적 리소스와 구성 정보를 제공한다.

제안 방법

  • Google Scholar, MedPub, Web of Science 및 주요 학회(ACL, EMNLP, NAACL 등)에서 문헌을 조사한다.
  • 데이터 소스, 모델 아키텍처, 사전 학습/미세 조정 패러다임에 따라 생물의학 PLMs를 분류한다.
  • 정보 추출, 분류, 질의응답, 대화 등과 같은 다운스트림 생물의학 태스크와 대응하는 PLM 방법을 요약한다.
  • 생물의학 PLMs의 한계와 향후 경향을 논의하여 향후 연구를 안내한다.
Figure 1 . Overview of selected released Biomedical pre-trained language models. One can see a more detailed list in Sec. 3 . Note that there is a BERT-like language model embedded in the overall architecture of AlphaFold 2.
Figure 1 . Overview of selected released Biomedical pre-trained language models. One can see a more detailed list in Sec. 3 . Note that there is a BERT-like language model embedded in the overall architecture of AlphaFold 2.

실험 결과

연구 질문

  • RQ1생물의학 PLMs를 사전 학습시키는 데 사용되는 데이터 소스(텍스트, 이미지, 단백질/DNA, 다중 모달리티)는 무엇인가?
  • RQ2도메인 적응 및 태스크 적응을 통해 일반 영역 PLMs에서 생물의학 PLMs가 어떻게 맞춤화되는가?
  • RQ3생물의학 PLMs의 주요 다운스트림 태스크와 벤치마크는 무엇이며 어떻게 다루어지는가?
  • RQ4생물의학 PLMs의 현재 한계는 무엇이며 어떤 향후 방향이 제안되는가?
  • RQ5생성적 모델 및 비전-언어 또는 단백질/DNA PLMs가 생물의학 연구 생태계에 어떻게 맞춰지는가?

주요 결과

  • 생물의학 PLMs은 텍스트를 넘어 비전-언어 모델 및 시퀀스 데이터(단백질/DNA)까지 확장된다.
  • 데이터 소스, 모델 아키텍처, 사전 학습 전략으로 PLMs를 분류하는 분류 체계가 제안된다.
  • 초보자들의 채택을 촉진하기 위해 리소스, 구성, 데이터셋, 대회, 장소를 하나로 정리한다.
  • 생물의학 PLMs는 한계와 도전에 직면해 있으며 향후 연구 방향과 도메인 적응을 촉진한다.
  • 비전-언어 및 단백질/DNA PLMs를 생물의학에서 다루고 다중 규모의 개관을 제공하는 최초의 고찰 중 하나이다.
Figure 2 . Parallel development of general and biomedical pre-trained language models. The time is determined by the released date of the paper, for example, in arXiv. General pre-trained language models are shown below in the timeline, and biomedical pre-trained language models are shown above the
Figure 2 . Parallel development of general and biomedical pre-trained language models. The time is determined by the released date of the paper, for example, in arXiv. General pre-trained language models are shown below in the timeline, and biomedical pre-trained language models are shown above the

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.