[論文レビュー] Pre-trained Language Models in Biomedical Domain: A Systematic Survey
このサーベイは生物医療分野の事前学習済み言語モデル(PLMs)を網羅的にレビューし、分類法を提案し、データソース・アーキテクチャ・タスク・ベンチマークを要約し、限界と今後の動向を強調します。テキスト、ビジョン-言語、タンパク質/ DNAのPLMsを含む生体医療分野。
Pre-trained language models (PLMs) have been the de facto paradigm for most natural language processing (NLP) tasks. This also benefits biomedical domain: researchers from informatics, medicine, and computer science (CS) communities propose various PLMs trained on biomedical datasets, e.g., biomedical text, electronic health records, protein, and DNA sequences for various biomedical tasks. However, the cross-discipline characteristics of biomedical PLMs hinder their spreading among communities; some existing works are isolated from each other without comprehensive comparison and discussions. It expects a survey that not only systematically reviews recent advances of biomedical PLMs and their applications but also standardizes terminology and benchmarks. In this paper, we summarize the recent progress of pre-trained language models in the biomedical domain and their applications in biomedical downstream tasks. Particularly, we discuss the motivations and propose a taxonomy of existing biomedical PLMs. Their applications in biomedical downstream tasks are exhaustively discussed. At last, we illustrate various limitations and future trends, which we hope can provide inspiration for the future research of the research community.
研究の動機と目的
- 限られた注釈データと豊富な未ラベルデータのため、生物医学ドメインでPLMsの利用を動機づける。
- データソース、アーキテクチャ、トレーニングパラダイムを横断して生物医学PLMsの全体像を要約する。
- 生物医学PLMsの分類法を提案し、研究者のための実用的なリソースと設定を提供する。
提案手法
- Google Scholar、MedPub、Web of Science、主要会議(ACL、EMNLP、NAACL など)から文献を調査する。
- データソース、モデルアーキテクチャ、事前学習/ファインチューニングパラダイムで生物医学PLMsを分類する。
- 下流の生物医学タスク(情報抽出、分類、QA、対話など)と対応するPLM手法を要約する。
- 限界と今後の動向を議論し、生物医学PLMsの今後の研究をガイドする。

実験結果
リサーチクエスチョン
- RQ1生物医学PLMsを事前学習させるために使用されるデータソース(テキスト、画像、タンパク質/ DNA、多模態)は何か?
- RQ2ドメイン適応とタスク適応によって一般ドメインのPLMsから生物医学PLMsはどう調整されるのか?
- RQ3生物医学PLMsの主な下流タスクとベンチマークは何で、どのように対処されているのか?
- RQ4生物医学PLMsの現在の限界は何で、今後の方向性は何か?
- RQ5生成モデルやビジョン-言語、タンパク質/ DNA PLMsは生物医学研究エコシステムにどのように適合するのか?
主な発見
- 生物医学PLMsはテキストを超えてビジョン-言語モデルや配列データ(タンパク質/ DNA)へと拡張している。
- データソース、モデルアーキテクチャ、事前学習戦略によってPLMsを分類する提案された分類法がある。
- 調査はリソース、設定、データセット、競技会、会場を統合し初心者の採用を促進する。
- 生物医学PLMsには限界と課題があり、今後の研究方向とドメイン適応を促す。
- ビジョン-言語とタンパク質/ DNA PLMsを生物医学で議論し、広くマルチスケールの概要を提供する初期の調査の一つである。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。