Skip to main content
QUICK REVIEW

[논문 리뷰] CancerLLM: A Large Language Model in Cancer Domain

Mingchen Li, Huang, Jiatan|arXiv (Cornell University)|2024. 06. 15.
Biomedical Text Mining and Ontologies인용 수 11
한 줄 요약

CancerLLM은 대형 임상 및 병리 데이터로 학습된 7B 암 도메인 LLM으로, 표현형 추출, 진단 생성, 치료 계획 생성에 대한 미세조정이 이루어졌으며, 최첨단 결과와 표적 테스트베드에서의 견고한 성능을 달성합니다.

ABSTRACT

Medical Large Language Models (LLMs) have demonstrated impressive performance on a wide variety of medical NLP tasks; however, there still lacks a LLM specifically designed for phenotyping identification and diagnosis in cancer domain. Moreover, these LLMs typically have several billions of parameters, making them computationally expensive for healthcare systems. Thus, in this study, we propose CancerLLM, a model with 7 billion parameters and a Mistral-style architecture, pre-trained on nearly 2.7M clinical notes and over 515K pathology reports covering 17 cancer types, followed by fine-tuning on two cancer-relevant tasks, including cancer phenotypes extraction and cancer diagnosis generation. Our evaluation demonstrated that the CancerLLM achieves state-of-the-art results with F1 score of 91.78% on phenotyping extraction and 86.81% on disganois generation. It outperformed existing LLMs, with an average F1 score improvement of 9.23%. Additionally, the CancerLLM demonstrated its efficiency on time and GPU usage, and robustness comparing with other LLMs. We demonstrated that CancerLLM can potentially provide an effective and robust solution to advance clinical research and practice in cancer domain

연구 동기 및 목표

  • 암 분야의 임상 NLP 작업을 개선하기 위해 암 특화 LLM의 개발을 촉진한다.
  • 암 데이터에 맞춘 Mistral 스타일 아키텍처의 7B 모델을 개발한다.
  • 표현형 추출, 진단 생성, 치료 계획 생성을 위한 세 가지 미세조정 데이터세트를 생성하고 활용한다.
  • 다양한 기준선 대비 생성 품질을 평가하고 반사적 사례와 오탈자에서의 강건성을 평가한다.

제안 방법

  • 2,676,642건의 암 임상 노트와 515,524건의 병리 보고서(총 17개 암 타입)에 대해 7B Mistral 스타일 LLM을 사전 학습한다.
  • LoRA 기반의 계속된 사전 학습을 적용하여 암 지식을 주입하고 특정 하이퍼파라미터(랭크 8, 알파 16, 드롭아웃 0.05, LR 2e-4)를 사용한다.
  • LoRA를 사용하여 랭크 64 및 알파 16으로 암 중심의 세 가지 작업에 대한 지시문 미세조정을 수행한다.
  • 비겹치지 않는 학습/테스트 분할을 가진 표현형 추출, 진단 생성, 치료 계획 생성을 위한 세 가지 하위 데이터세트를 구성한다.
  • 정확 매치(Exact Match), BLEU-2, ROUGE-L 지표로 평가하고 반사적 사례와 오탈자 같은 강건성 테스트베드를 포함한다.
  • 7B, 8B, 13B, 70B 모델에 걸친 14개 기준선과 비교하고 생성 품질과 효율성(시간 및 GPU 메모리)을 모두 보고한다.
Figure 1: The evolution of medical LLM performance on three tasks—cancer phenotype extraction, diagnosis generation, and treatment plan generation—is measured using the average F1 score, which includes Exact Match, BLEU-2, and ROUGE-L. Our CancerLLM achieves the highest performance with an F1 score
Figure 1: The evolution of medical LLM performance on three tasks—cancer phenotype extraction, diagnosis generation, and treatment plan generation—is measured using the average F1 score, which includes Exact Match, BLEU-2, and ROUGE-L. Our CancerLLM achieves the highest performance with an F1 score

실험 결과

연구 질문

  • RQ17B 암 도메인 LLM이 암 표현형 추출, 진단 생성, 치료 계획 생성에서 최첨단 생성 품질을 달성할 수 있는가?
  • RQ2연속 사전 학습 및 지시문 미세조정을 통한 도메인 특정 암 지식 주입이 더 큰 일반 의학 LLM보다 성능이 더 좋은가?
  • RQ3임상 텍스트에서 반사적 라벨과 오탈자에 대해 CancerLLM의 강건성은 어느 정도인가?
  • RQ4임상 현장에서 소형 암 도메인 LLM을 배치할 때 생성 시간과 메모리 사용의 트레이드오프는 무엇인가?

주요 결과

  • 평가된 모델 중 세 가지 작업에서 전반적으로 최상의 성능을 달성했으며, 진단 생성에서 기준선 대비 평균 F1이 8.1% 증가했다.
  • 암 진단 생성에서 CancerLLM은 평균 F1 86.81 및 EM 83.50을 달성하며 7B, 13B, 70B 기준선을 모두 능가했다.
  • 암 치료 계획 생성에서 CancerLLM은 평균 F1 91.78 및 EM 89.37을 달성하며 다시 getest된 모델들 가운데 선두를 차지했다.
  • 암 표현형 추출에서 CancerLLM은 평균 F1 93.98 및 EM 89.37로 대형 모델에 근접하거나 이를 상회하는 성능을 훨씬 적은 파라미터로 달성했다.
  • 강건성 테스트베드는 반사적 섭동과 오탈자 하에서도 CancerLLM이 경쟁력 있는 성능을 유지하는 것을 보여주며, 섭동 비율이 증가할수록 일부 저하가 발생하나 노이즈가 높은 경우에도 강력한 기준선보다 종종 더 우수하다(예: 반사적 비율이 80%인 경우).
  • CancerLLM은 추론 시간 1:14:12 및 표현형 추출에 대한 GPU 메모리 사용 5,550 MB로 우수한 효율성을 보여주며 여러 70B 대비 현저히 낮은 편이다.
Figure 2: Overview of CancerLLM
Figure 2: Overview of CancerLLM

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.