[논문 리뷰] A Comparative Study of Open-Source Large Language Models, GPT-4 and Claude 2: Multiple-Choice Test Taking in Nephrology
이 논문은 신장학 MCQ에서 오픈 소스 LLM을 GPT-4와 Claude 2와 비교하여 GPT-4와 Claude 2가 오픈 소스 모델보다 큰 차이로 우수하다고 나타냈다.
In recent years, there have been significant breakthroughs in the field of natural language processing, particularly with the development of large language models (LLMs). These LLMs have showcased remarkable capabilities on various benchmarks. In the healthcare field, the exact role LLMs and other future AI models will play remains unclear. There is a potential for these models in the future to be used as part of adaptive physician training, medical co-pilot applications, and digital patient interaction scenarios. The ability of AI models to participate in medical training and patient care will depend in part on their mastery of the knowledge content of specific medical fields. This study investigated the medical knowledge capability of LLMs, specifically in the context of internal medicine subspecialty multiple-choice test-taking ability. We compared the performance of several open-source LLMs (Koala 7B, Falcon 7B, Stable-Vicuna 13B, and Orca Mini 13B), to GPT-4 and Claude 2 on multiple-choice questions in the field of Nephrology. Nephrology was chosen as an example of a particularly conceptually complex subspecialty field within internal medicine. The study was conducted to evaluate the ability of LLM models to provide correct answers to nephSAP (Nephrology Self-Assessment Program) multiple-choice questions. The overall success of open-sourced LLMs in answering the 858 nephSAP multiple-choice questions correctly was 17.1% - 25.5%. In contrast, Claude 2 answered 54.4% of the questions correctly, whereas GPT-4 achieved a score of 73.3%. We show that current widely used open-sourced LLMs do poorly in their ability for zero-shot reasoning when compared to GPT-4 and Claude 2. The findings of this study potentially have significant implications for the future of subspecialty medical training and patient care.
연구 동기 및 목표
- nephSAP MCQs에서 오픈 소스 LLM의 신장학 지식 역량 평가.
- 신장학 데이터셋에서 제로샷 성능으로 오픈 소스 LLM과 GPT-4 및 Claude 2를 비교.
- 모델 설명의 품질 및 실제 정답과의 의미론적 정렬 분석.
제안 방법
- 858개의 nephSAP MCQ에서 Koala 7B, Falcon 7B, Stable-Vicuna 13B, 및 Orca Mini 13B를 GPT-4 및 Claude 2와 비교 평가.
- 오픈 소스 LLM에서 순전파를 위한 Context, Question, Choices를 연결하여 입력 프롬트를 구성.
- 정규 표현식 기반 추출로 자동 출력을 구문 분석하여 예측된 답안을 결정하고 정답과 비교.
- 정확도는 원시 정답 개수로 측정하고 신장학 하위 주제별로 주제별 성능을 계산.
- 정답 설명과 ground-truth 설명에 대해 BLEU, WER, 코사인 유사도를 사용하여 설명을 평가.
실험 결과
연구 질문
- RQ1open-source LLM이 nephSAP 신장학 MCQ에서 GPT-4 및 Claude 2와 비교하여 어떤 성능을 보이는가?
- RQ2신장학에서 오픈 소스 LLM의 주제별 강점과 약점은 무엇인가?
- RQ3오픈 소스 LLM의 설명 품질이 정답과 비교해 어떤가?
- RQ4오픈 소스 LLM과 GPT-4/Claude 2 간의 성능 격차를 설명할 수 있는 요인은 무엇인가?
주요 결과
- GPT-4는 629개 정답(73.3%), Claude 2는 467개(54.4%), 그리고 오픈 소스 LLM은 17.1%에서 25.5% 정답으로 범위.
- 오픈 소스 모델 중 Vicuna가 219 정답(25.5%)으로 최고를 기록했고, Koala가 204정답(23.8%)으로 그 뒤를 이었다.
- Falcon 155(18.1%), Orca-Mini 147(17.1%), Koala 204(23.8%).
- 질문 구조상 무작위 추측은 23.8%를 보일 것이므로, 오픈 소스 모델 중 Koala만이 우연 기대치를 약간 상회했다.
- 오픈 소스 모델의 설명은 BLEU 점수가 낮게 나타났고(예: Vicuna 9%, Falcon 8%, Orca-Mini 5%, Koala 5%), 코사인 유사도 점수도 모든 모델에서 보조적이지 않았다.
- GPT-4와 Claude 2는 모든 신장학 주제에서 오픈 소스 모델을 능가했으며, GPT-4는 대부분의 주제에서 인간에 가까운 성능을 보였다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.