[Paper Review] Clinical Validation of Medical-based Large Language Model Chatbots on Ophthalmic Patient Queries with LLM-based Evaluation
The paper empirically evaluates four small medical LLMs on ophthalmic patient queries and compares LLM-based evaluation with clinician grading to assess safety, consensus, and clinical depth.
Domain specific large language models are increasingly used to support patient education, triage, and clinical decision making in ophthalmology, making rigorous evaluation essential to ensure safety and accuracy. This study evaluated four small medical LLMs Meerkat-7B, BioMistral-7B, OpenBioLLM-8B, and MedLLaMA3-v20 in answering ophthalmology related patient queries and assessed the feasibility of LLM based evaluation against clinician grading. In this cross sectional study, 180 ophthalmology patient queries were answered by each model, generating 2160 responses. Models were selected for parameter sizes under 10 billion to enable resource efficient deployment. Responses were evaluated by three ophthalmologists of differing seniority and by GPT-4-Turbo using the S.C.O.R.E. framework assessing safety, consensus and context, objectivity, reproducibility, and explainability, with ratings assigned on a five point Likert scale. Agreement between LLM and clinician grading was assessed using Spearman rank correlation, Kendall tau statistics, and kernel density estimate analyses. Meerkat-7B achieved the highest performance with mean scores of 3.44 from Senior Consultants, 4.08 from Consultants, and 4.18 from Residents. MedLLaMA3-v20 performed poorest, with 25.5 percent of responses containing hallucinations or clinically misleading content, including fabricated terminology. GPT-4-Turbo grading showed strong alignment with clinician assessments overall, with Spearman rho of 0.80 and Kendall tau of 0.67, though Senior Consultants graded more conservatively. Overall, medical LLMs demonstrated potential for safe ophthalmic question answering, but gaps remained in clinical depth and consensus, supporting the feasibility of LLM based evaluation for large scale benchmarking and the need for hybrid automated and clinician review frameworks to guide safe clinical deployment.
Motivation & Objective
- Assess safety, accuracy, and clinical usefulness of domain-specific LLM chatbots in ophthalmology.
- Evaluate feasibility of LLM-based evaluation as a benchmarking tool against clinician assessments.
- Identify gaps in clinical depth, consensus, and potential hallucinations in model responses.
- Explore resource-efficient deployment by selecting models under 10B parameters.
Proposed method
- Answer 180 ophthalmology patient queries with each model to generate 2160 responses.
- Have three ophthalmologists of differing seniority and GPT-4-Turbo grade responses using the S.C.O.R.E. framework on a 5-point Likert scale.
- Use Spearman rank correlation, Kendall tau, and kernel density estimates to compare LLM vs clinician grading.
- Select models with parameter sizes under 10 billion for resource-efficient deployment.
Experimental results
Research questions
- RQ1Can small medical LLMs provide safe and clinically useful ophthalmic answers?
- RQ2How well does LLM-based evaluation align with clinician grading in ophthalmology queries?
- RQ3What are the rates of hallucinations or clinically misleading content across models?
- RQ4Is LLM-based benchmarking feasible for large-scale evaluation of medical chatbots in ophthalmology?
Key findings
- Meerkat-7B achieved the highest mean scores across Senior Consultants, Consultants, and Residents (3.44, 4.08, 4.18 respectively).
- MedLLaMA3-v20 had the poorest performance with 25.5% of responses containing hallucinations or clinically misleading content.
- GPT-4-Turbo grading showed strong alignment with clinician assessments (Spearman 0.80, Kendall 0.67) but conservative grading by Senior Consultants.
- Medical LLMs show potential for safe ophthalmic QA but gaps remain in clinical depth and consensus.
- LLM-based evaluation appears feasible for large-scale benchmarking and supports a hybrid automated+clinician review framework.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.