Skip to main content
QUICK REVIEW

[논문 리뷰] Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation

Shefayat E Shams Adib, Ahmed Alfey Sani|arXiv (Cornell University)|2026. 02. 16.
Artificial Intelligence in Healthcare and Education인용 수 0
한 줄 요약

이 논문은 iCliniq 데이터를 사용한 제로샷 평가에서 다섯 개 LLM을 의료 Q&A에 대해 벤치마킹하고, 자동 지표(BLEU/ROUGE)와 LLM-판단자 평가를 비교하여 의료 정확도와 안전성을 측정한다. 더 큰 모델이 더 우수한 성능을 보이며, Llama 3.3 70B Instruct가 선두를 이끈다.

ABSTRACT

Recently, Large Language Models (LLMs) have gained significant traction in medical domain, especially in developing a QA systems to Medical QA systems for enhancing access to healthcare in low-resourced settings. This paper compares five LLMs deployed between April 2024 and August 2025 for medical QA, using the iCliniq dataset, containing 38,000 medical questions and answers of diverse specialties. Our models include Llama-3-8B-Instruct, Llama 3.2 3B, Llama 3.3 70B Instruct, Llama-4-Maverick-17B-128E-Instruct, and GPT-5-mini. We are using a zero-shot evaluation methodology and using BLEU and ROUGE metrics to evaluate performance without specialized fine-tuning. Our results show that larger models like Llama 3.3 70B Instruct outperform smaller models, consistent with observed scaling benefits in clinical tasks. It is notable that, Llama-4-Maverick-17B exhibited more competitive results, thus highlighting evasion efficiency trade-offs relevant for practical deployment. These findings align with advancements in LLM capabilities toward professional-level medical reasoning and reflect the increasing feasibility of LLM-supported QA systems in the real clinical environments. This benchmark aims to serve as a standardized setting for future study to minimize model size, computational resources and to maximize clinical utility in medical NLP applications.

연구 동기 및 목표

  • 현대 LLM의 의료 Q&A에 대한 대규모 실제 데이터셋(iCliniq)에서의 포괄적 제로샷 벤치마크 제공
  • 모델 크기/아키텍처와 의료 Q&A 성능 간의 상관관계 평가
  • 자동 지표와 LLM-판단자 임상 품질 평가를 결합한 표준화된 이중 평가 프레임워크의 도입 및 검증
  • 임상 환경의 정확도와 자원 제약 간의 균형을 고려한 배포 가이드 제공

제안 방법

  • 다섯 개 LLM에 대해 표준화된 의료 프롬프트를 활용한 제로샷 평가 프로토콜 사용
  • 38,000개의 iCliniq Medical QA 데이터셋 중 3,000문항 서브셋으로 평가
  • 어휘 유사도 및 커버리지를 평가하기 위한 BLEU 및 ROUGE 지표 계산
  • 5점 척도(가중점수 30/25/20/15/10)로 Medical Accuracy, Completeness, Safety, Clarity, Helpfulness를 평가하는 LLM-판단자 프레임워크(Claude Sonnet 4) 적용
  • 향상된 부분을 맥락화하기 위해 prior 연구의 MedLM 벤치마크와 결과 비교

실험 결과

연구 질문

  • RQ1iCliniq 데이터셋을 사용한 다섯 가지 현대 LLM의 제로샷 의료 QA 태스크 성능은 어떠한가?
  • RQ2제로샷 설정에서 모델 크기/아키텍처와 의료 QA 성능 간의 관계는 무엇인가?
  • RQ3의료 QA에서 LLM-판단자 평가가 전통적 BLEU/ROUGE 지표와 어떻게 정렬되는가?
  • RQ4높은 정확도 임상 환경과 자원 제약 설정 간의 배포 가이드는 무엇을 제시하는가?

주요 결과

  • Llama 3.3 70B Instruct가 평가된 모델들 중에서 BLEU-1, ROUGE-1, ROUGE-L에서 가장 높은 성능을 달성했다.
  • Llama-4-Maverick 17B는 파라미터 수가 훨씬 적으면서도 70B 모델에 근접한 성능으로 효율성이 경쟁력 있다.
  • GPT-5-mini는 자동 지표에서 전반적으로 낮은 성능을 보였으며 구현/구성 이슈를 시사한다.
  • 모델 크기와 의료 QA 성능 간 명확한 양의 상관관계가 있으며, 아키텍처 혁신으로 작은 모델이 더 큰 모델에 근접할 수 있다.
  • LLM-판단자 결과는 자동 지표와 일치하여 순위를 확정하고 평가 프레임워크를 입증한다.
  • 의료 정확도는 상위 모델에서 가장 높고(최상 4.83/5), 안전성은 GPT-5-mini에서 가장 높게 나타나지만 어휘 지표는 다소 약하다(3.80/5).

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.