[논문 리뷰] Biomedical Large Languages Models Seem not to be Superior to Generalist Models on Unseen Medical Data
생물의학 LLM은 도메인 데이터에 대해 미세조정하더라도 보지 못한 의료 데이터에 대한 여러 임상 과제에서 일반 목적 모델에 비해 일반적으로 성능이 떨어지는 경향이 있어, 생물의학 미세조정의 한계가 시사되며 검색 보강 접근법을 유망한 대안으로 강조한다.
Large language models (LLMs) have shown potential in biomedical applications, leading to efforts to fine-tune them on domain-specific data. However, the effectiveness of this approach remains unclear. This study evaluates the performance of biomedically fine-tuned LLMs against their general-purpose counterparts on a variety of clinical tasks. We evaluated their performance on clinical case challenges from the New England Journal of Medicine (NEJM) and the Journal of the American Medical Association (JAMA) and on several clinical tasks (e.g., information extraction, document summarization, and clinical coding). Using benchmarks specifically chosen to be likely outside the fine-tuning datasets of biomedical models, we found that biomedical LLMs mostly perform inferior to their general-purpose counterparts, especially on tasks not focused on medical knowledge. While larger models showed similar performance on case tasks (e.g., OpenBioLLM-70B: 66.4% vs. Llama-3-70B-Instruct: 65% on JAMA cases), smaller biomedical models showed more pronounced underperformance (e.g., OpenBioLLM-8B: 30% vs. Llama-3-8B-Instruct: 64.3% on NEJM cases). Similar trends were observed across the CLUE (Clinical Language Understanding Evaluation) benchmark tasks, with general-purpose models often performing better on text generation, question answering, and coding tasks. Our results suggest that fine-tuning LLMs to biomedical data may not provide the expected benefits and may potentially lead to reduced performance, challenging prevailing assumptions about domain-specific adaptation of LLMs and highlighting the need for more rigorous evaluation frameworks in healthcare AI. Alternative approaches, such as retrieval-augmented generation, may be more effective in enhancing the biomedical capabilities of LLMs without compromising their general knowledge.
연구 동기 및 목표
- 생물이 아닌 임상의 unseen 데이터와 과제에서 생물의학 미세조정이 LLM 성능을 향상시키는지 평가한다.
- 다양한 임상 벤치마크에서 생물의학 미세조정 LLM과 일반 목적형 기준 모델을 비교한다.
- 도메인 적응이 의료 AI에서 이익을 제공하는 영역과 그렇지 않은 영역을 조사한다.
제안 방법
- NEJM 및 JAMA 사례 도전 과제에서 생물의학 및 일반-purpose LLM을 평가한다(347 NEJM, 140 JAMA 문제).
- CLUE의 MeDiSumQA, MeDiSumCode, MedNLI, MeQSum, ProblemSummary 및 LongHealth 벤치마크를 평가한다.
- 표준화된 프롬프트 및 추론 설정을 통해 작업별 고정 평가 지표(정확도, F1, ROUGE, BERTScore)를 사용한다.
- 다양한 크기와 아키텍처(Llama, Mistral, OpenBioLLM 등)와 그들의 채팅/지시 변형을 포함한다.
- 데이터 누출을 피하기 위해 벤치마크가 생물의학 미세조정 데이터 외부일 가능성이 있도록 한다.

실험 결과
연구 질문
- RQ1생물이 아닌 미세조정 LLM이 unseen 임상 사례 데이터에서 일반 목적 LLM보다 성능이 우수한가?
- RQ2도메인 특화 LLM이 정보추출, 코딩, 요약 과제에서 일반 모델과 비교하여 어떤 성능을 보이는가?
- RQ3일반 목적 모델의 우위가 길이가 긴 임상 문서와 환각이 잦은 작업에서 일관되는가?
- RQ4의료 LLM에서 도메인 특화 미세조정보다 검색 보강 생성 방식이 더 효과적일 수 있는가?
주요 결과
- JAMA 및 NEJM 사례 도전 과제에서 다수의 일반 모델(OpenBioLLM-70B, Llama-3-70B-Instruct 등)이 최고 정확도(예: 66-74%)를 달성했다.
- Llama-3-8B-Instruct는 생물의학 모델보다 자주 더 우수한 성능을 보이는 경우가 많았으며(NEJM에서 64% 대 18%, JAMA에서 64% 대 18% 등), 종종 생물의학 모델을 능가했다.
- MedNLI, ProblemSummary 및 MeQSum 전반에서 모든 생물의학 LLM은 일반 모델에 비해 성능이 낮았다.
- MeDiSumCode 및 일부 장기 건강 관련 과제에서는 심층 지식 과제와 긴 맥락 처리에서 일반 모델의 강점이 더 강조되었다.
- LongHealth 과제 결과에서 생물의학 모델은 더 많은 환각 현상을 보였고 일반 모델이 환각 관련 평가에서 상대적으로 더 우수했다.
- 전반적으로 더 큰 모델일수록 생물의학 모델과 일반 모델 간의 성능 차이가 작아져, 미세조정만으로는 도메인 적응에 충분하지 않을 수 있음을 시사한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.