Skip to main content
QUICK REVIEW

[논문 리뷰] Med-BERT: pre-trained contextualized embeddings on large-scale structured electronic health records for disease prediction

Laila Rasmy, Yang Xiang|arXiv (Cornell University)|2020. 05. 22.
Machine Learning in Healthcare참고 문헌 43인용 수 61
한 줄 요약

Med-BERT는 BERT 프레임워크를 대규모 구조화된 EHR 데이터에 적용하여 맥락화된 임베딩을 생성하고, 특히 작은 파인튜닝 세트에서 질병 예측 성능을 향상시킨다.

ABSTRACT

Deep learning (DL) based predictive models from electronic health records (EHR) deliver impressive performance in many clinical tasks. Large training cohorts, however, are often required to achieve high accuracy, hindering the adoption of DL-based models in scenarios with limited training data size. Recently, bidirectional encoder representations from transformers (BERT) and related models have achieved tremendous successes in the natural language processing domain. The pre-training of BERT on a very large training corpus generates contextualized embeddings that can boost the performance of models trained on smaller datasets. We propose Med-BERT, which adapts the BERT framework for pre-training contextualized embedding models on structured diagnosis data from 28,490,650 patients EHR dataset. Fine-tuning experiments are conducted on two disease-prediction tasks: (1) prediction of heart failure in patients with diabetes and (2) prediction of pancreatic cancer from two clinical databases. Med-BERT substantially improves prediction accuracy, boosting the area under receiver operating characteristics curve (AUC) by 2.02-7.12%. In particular, pre-trained Med-BERT substantially improves the performance of tasks with very small fine-tuning training sets (300-500 samples) boosting the AUC by more than 20% or equivalent to the AUC of 10 times larger training set. We believe that Med-BERT will benefit disease-prediction studies with small local training datasets, reduce data collection expenses, and accelerate the pace of artificial intelligence aided healthcare.

연구 동기 및 목표

  • 레이블 데이터가 제한될 때 구조화된 EHR 데이터에 대해 사전 학습된 맥락화 임베딩의 사용을 고무하여 질병 예측을 향상시키는 것을 목표로 한다.
  • 대규모 EHR 말뭉치를 활용하여 아래 임상 예측 작업에 전이될 수 있는 표현을 사전 학습한다.
  • 데이터가 드문 설정에서 특정 질병에 대한 예측 정확도의 향상을 입증한다.

제안 방법

  • 매우 대규모 EHR 코호트(28,490,650명)의 구조화된 진단 데이터에 BERT 프레임워크를 적용한다.
  • 구조화된 EHR 데이터에서 맥락화된 임베딩을 사전 학습하고 다운스트림 작업에서 파인튜닝한다.
  • 두 개의 임상 데이터베이스에서 당뇨병 환자의 심부전 및 췌장암이라는 두 가지 질병 예측 작업에 대해 평가한다.
  • 사전 학습이 없는 기준선과의 성능을 비교하여 사전 학습의 이득을 정량화한다.
  • 다양한 파인튜닝 데이터 크기에 따른 AUC 향상을 보고하고, 작은 데이터의 이점을 강조한다.

실험 결과

연구 질문

  • RQ1구조화된 EHR 데이터에서 사전 학습된 Med-BERT가 비사전 학습 모델과 비교하여 질병 예측 성능을 향상시킬 수 있는가?
  • RQ2파인튜닝 데이터가 희소할 때(예: 300–500 샘플) 사전 학습이 성능에 어떤 영향을 미치는가?
  • RQ3다른 질병과 임상 데이터베이스 전반에 걸쳐 이득이 일반화되는가?

주요 결과

  • Med-BERT는 예측 정확도를 크게 향상시켜, 각 작업에서 AUC를 2.02–7.12% 증가시킨다.
  • 사전 학습된 Med-BERT는 매우 작은 파인튜닝 세트(300–500 샘플)에서 성능을 크게 향상시켜 AUC를 20% 이상 증가시킨다.
  • 작은 데이터에서의 이득은 10배 더 큰 훈련 세트로 달성될 수 있는 것과 비슷하다.
  • 이 방법은 작은 지역 데이터 세트로 질병 예측 연구를 지원하고 데이터 수집 비용을 감소시킬 수 있다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.