[논문 리뷰] Pre-trained Language Models in Biomedical Domain: A Survey from Multiscale Perspective
이 종합 검토는 생물의학 분야의 사전 훈련된 언어 모델(PLM)에 대해 체계적이고 다스케일적인 리뷰를 제공하며, 기존 모델을 분류하고 용어와 벤치마크를 표준화하며, 하류 작업에서의 응용을 분석한다. 이는 통합된 분류 체계를 제공하고, 교차 분야 발전을 위한 핵심적 한계와 향후 연구 방향을 규명한다.
Pre-trained language models have been the de facto paradigm for most natural language processing (NLP) tasks. In the biomedical domain, which also benefits from NLP techniques, various pre-trained language models were proposed by leveraging domain datasets including biomedical literature, biomedical social medial, electronic health records, and other biological sequences. Large amounts of efforts have been explored on applying these biomedical pre-trained language models to downstream biomedical tasks, from informatics, medicine, and computer science (CS) communities. However, it seems that the vast majority of existing works are isolated from each other probably because of the cross-discipline characteristics. It is expected to propose a survey that not only systematically reviews recent advances of biomedical pre-trained language models and their applications but also standardizes terminology, taxonomy, and benchmarks. Therefore, this paper summarizes the recent progress of pre-trained language models used in the biomedical domain. Particularly, an overview and taxonomy of existing biomedical pre-trained language models as well as their applications in biomedical downstream tasks are exhaustively discussed. At last, we illustrate various limitations and future trends, which we hope can provide inspiration for the future research.
연구 동기 및 목표
- 다양한 척도에서 최근 생물의학 분야의 사전 훈련된 언어 모델의 발전을 체계적으로 검토하기 위해.
- 생물의학 PLM에 대한 용어, 분류 체계, 평가 벤치마크를 표준화하기 위해.
- 생물의학, 인포매틱스, 컴퓨터 과학 분야의 고립된 연구 노력 간 격차를 메우기 위해.
- 생물의학 NLP 분야에서의 한계와 새로운 추세를 규명하기 위해 향후 연구를 위한 기반을 마련하기 위해.
제안 방법
- 생물의학 문헌, EHR, 생물학적 서열과 같은 훈련 데이터 소스에 기반한 생물의학 PLM의 분류.
- 아키텍처, 사전 훈련 목표, 영역 특이성에 따라 모델을 분류하기 위한 통합된 분류 체계 구축.
- 이름 있는 실체 인식, 관계 추출, 질의 응답과 같은 하류 생물의학 작업에서의 모델 응용 분석.
- 재현 가능성과 비교 가능성을 보장하기 위해 기존의 벤치마크와 평가 프로토콜을 통합.
- 모델 일반화 및 영역 적응 문제에서 반복적으로 나타나는 방법론적 격차와 과제 규명.
실험 결과
연구 질문
- RQ1생물의학 사전 훈련된 언어 모델은 아키텍처, 사전 훈련 데이터, 목적으로 어떻게 분류되는가?
- RQ2기존 생물의학 PLM 간 설계 및 응용 측면에서의 주요 차이점과 유사점은 무엇인가?
- RQ3질병 예측 및 약물 반응 모델링과 같은 다양한 하류 작업에서 생물의학 PLM의 성능은 어떻게 되는가?
- RQ4생물의학 NLP 연구 전반에서 용어, 벤치마크, 평가 방법론의 표준화 과제는 무엇인가?
- RQ5생물의학 PLM 개발 분야에서의 주요 한계와 새로운 연구 추세는 무엇인가?
주요 결과
- 데이터 소스, 아키텍처, 사전 훈련 목표를 포함한 생물의학 사전 훈련된 언어 모델의 종합적 분류 체계가 수립되었다.
- 생물의학 NLP 연구 전반에서 용어, 벤치마크, 평가 프로토콜의 표준화 부족이 규명되었다.
- 기존 모델은 도메인 특화 데이터로 미세 조정된 경우 이름 있는 실체 인식 및 관계 추출과 같은 작업에서 뛰어난 성능을 보였다.
- 진전이 있었음에도 불구하고, 생물의학, 인포매틱스, 컴퓨터 과학 분야의 고립된 연구 노력은 교차 분야 발전을 저해한다.
- 향후 연구는 모델 일반화 향상, 데이터 편향 감소, 재현 가능성을 위한 공동 벤치마크 개발에 초점을 맞춰야 한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.