Skip to main content
QUICK REVIEW

[논문 리뷰] Clinical Validation of Medical-based Large Language Model Chatbots on Ophthalmic Patient Queries with LLM-based Evaluation

Ting Fang Tan, Kabilan Elangovan|arXiv (Cornell University)|2026. 02. 05.
Artificial Intelligence in Healthcare and Education인용 수 0
한 줄 요약

이 논문은 네 가지 소형 의료 LLM을 안과 환자 질의에 대해 경험적으로 평가하고 LLM 기반 평가를 임상의 등급과 비교하여 안전성, 합의 및 임상적 심도를 평가한다.

ABSTRACT

Domain specific large language models are increasingly used to support patient education, triage, and clinical decision making in ophthalmology, making rigorous evaluation essential to ensure safety and accuracy. This study evaluated four small medical LLMs Meerkat-7B, BioMistral-7B, OpenBioLLM-8B, and MedLLaMA3-v20 in answering ophthalmology related patient queries and assessed the feasibility of LLM based evaluation against clinician grading. In this cross sectional study, 180 ophthalmology patient queries were answered by each model, generating 2160 responses. Models were selected for parameter sizes under 10 billion to enable resource efficient deployment. Responses were evaluated by three ophthalmologists of differing seniority and by GPT-4-Turbo using the S.C.O.R.E. framework assessing safety, consensus and context, objectivity, reproducibility, and explainability, with ratings assigned on a five point Likert scale. Agreement between LLM and clinician grading was assessed using Spearman rank correlation, Kendall tau statistics, and kernel density estimate analyses. Meerkat-7B achieved the highest performance with mean scores of 3.44 from Senior Consultants, 4.08 from Consultants, and 4.18 from Residents. MedLLaMA3-v20 performed poorest, with 25.5 percent of responses containing hallucinations or clinically misleading content, including fabricated terminology. GPT-4-Turbo grading showed strong alignment with clinician assessments overall, with Spearman rho of 0.80 and Kendall tau of 0.67, though Senior Consultants graded more conservatively. Overall, medical LLMs demonstrated potential for safe ophthalmic question answering, but gaps remained in clinical depth and consensus, supporting the feasibility of LLM based evaluation for large scale benchmarking and the need for hybrid automated and clinician review frameworks to guide safe clinical deployment.

연구 동기 및 목표

  • 안과 분야 특화 LLM 챗봇의 안전성, 정확성 및 임상적 유용성 평가.
  • 임상 평가와의 비교를 위한 벤치마킹 도구로서의 LLM 기반 평가의 실현 가능성 평가.
  • 모델 응답에서의 임상적 심도, 합의 및 잠재적 환각(hallucination) 여부의 격차 식별.
  • 매개변수가 10B 미만인 모델을 선별하여 자원 효율적인 배치를 모색.

제안 방법

  • 각 모델로 180건의 안과 환자 질의에 응답하여 총 2160개의 응답을 생성한다.
  • 다른 경력의 안과 전문의 3명과 GPT-4-Turbo가 S.C.O.R.E. 프레임워크를 사용하여 5점 리커트 척도에서 응답을 평가한다.
  • LLM 대 임상의 등급을 비교하기 위해 스피어먼 순위상관, 켄달 타우 및 커널 밀도 추정치를 사용한다.
  • 자원 효율적 배치를 위해 매개변수 크기가 10B 미만인 모델을 선택한다.

실험 결과

연구 질문

  • RQ1소형 의학 LLM이 안전하고 임상적으로 유용한 안과 답변을 제공할 수 있는가?
  • RQ2LLM 기반 평가가 안과 질의에서 임상의 등급과 얼마나 잘 일치하는가?
  • RQ3모델 간 환각 또는 임상적으로 오해를 불러일으킬 수 있는 콘텐츠의 비율은 어느 정도인가?
  • RQ4안과 분야의 의학 챗봇 대규모 평가를 위한 LLM 기반 벤치마킹이 실현 가능한가?

주요 결과

  • Meerkat-7B는 Senior Consultants, Consultants, 및 Residents에서 가장 높은 평균 점수를 달성했다 (3.44, 4.08, 4.18 각각).
  • MedLLaMA3-v20은 25.5%의 응답에서 환각이나 임상적으로 오해를 포함한 콘텐츠를 포함해 최하의 성능을 보였다.
  • GPT-4-Turbo 등급은 임상의 평가와 강한 일치를 보였고(Spearman 0.80, Kendall 0.67) 그러나 Senior Consultants에 의한 등급은 보수적이었다.
  • 의료 LLM은 안전한 안과 QA 가능성을 보이지만 임상 심도와 합의의 격차가 남아 있다.
  • LLM 기반 평가가 대규모 벤치마킹에 실현 가능해 보이며 자동화+임상의 리뷰 하이브리드 프레임워크를 지지한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.