[논문 리뷰] GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records
GatorTron은 대규모 임상 언어 모델(최대 8.9B 파라미터)을 개발하여 >90B 단어 이상으로 학습하고(그 중 >82B의 비식별 임상 텍스트 포함), 이를 5개의 임상 NLP 과제에서 평가하여 규모 확장에 따른 유의미한 성능 향상을 보여준다.
There is an increasing interest in developing artificial intelligence (AI) systems to process and interpret electronic health records (EHRs). Natural language processing (NLP) powered by pretrained language models is the key technology for medical AI systems utilizing clinical narratives. However, there are few clinical language models, the largest of which trained in the clinical domain is comparatively small at 110 million parameters (compared with billions of parameters in the general domain). It is not clear how large clinical language models with billions of parameters can help medical AI systems utilize unstructured EHRs. In this study, we develop from scratch a large clinical language model - GatorTron - using >90 billion words of text (including >82 billion words of de-identified clinical text) and systematically evaluate it on 5 clinical NLP tasks including clinical concept extraction, medical relation extraction, semantic textual similarity, natural language inference (NLI), and medical question answering (MQA). We examine how (1) scaling up the number of parameters and (2) scaling up the size of the training data could benefit these NLP tasks. GatorTron models scale up the clinical language model from 110 million to 8.9 billion parameters and improve 5 clinical NLP tasks (e.g., 9.6% and 9.5% improvement in accuracy for NLI and MQA), which can be applied to medical AI systems to improve healthcare delivery. The GatorTron models are publicly available at: https://catalog.ngc.nvidia.com/orgs/nvidia/teams/clara/models/gatortron_og.
연구 동기 및 목표
- 대 비정형 EHR 데이터를 보다 잘 활용하기 위해 대규모 임상 언어 모델 개발의 동기를 제시한다.
- 임계 파라미터 수와 데이터 규모의 확장이 임상 NLP 과제 성능에 미치는 영향을 조사한다.
- 다수의 임상 NLP 과제에서 GatorTron을 체계적으로 평가하여 과제 간 일반화를 평가한다.
제안 방법
- >90B 단어의 텍스트로부터 GatorTron을 처음부터 학습시키되 >82B의 비식별 임상 텍스트를 포함한다.
- 성능 향상을 연구하기 위해 모델을 110M에서 8.9B 파라미터로 확장한다.
- 5개의 임상 NLP 과제에서 평가한다: 임상 개념 추출, 의학 관계 추출, 의미적 텍스트 유사성, 자연어 추론(NLI), 의학 질문 응답(MQA).
- 파라미터 확장과 데이터 규모의 효과를 평가하기 위해 모델 규모 간의 성능 차이를 비교한다.
실험 결과
연구 질문
- RQ1모델 크기(파라미터)를 증가시키면 임상 NLP 과제의 성능에 어떤 영향을 미치는가?
- RQ2학습 데이터 크기를 증가시키면 과제 전반에 걸친 결과에 어떤 영향을 미치는가?
- RQ3대규모 임상 LMs가 추출, 관계, 유사성, NLI 및 QA와 같은 임상 NLP 과제에서 일관된 이점을 제공하는가?
- RQ4파라미터와 데이터를 함께 확장할 때 주요 과제에서의 상대적 개선은 무엇인가?
주요 결과
- 모델 규모를 110M에서 8.9B 파라미터로 확장하면 5개 임상 NLP 과제 전반의 성능이 향상된다.
- NLI 정확도는 확장으로 9.6% 개선된다.
- MQA 정확도는 확장으로 9.5% 개선된다.
- >90B 단어의 학습, 그 중 >82B의 비식별 임상 텍스트를 포함하여, 상당한 성능 향상을 뒷받침한다.
- GatorTron 모델은 의료 인공지능 시스템에서 공개적으로 이용 가능하다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.