[Paper Review] EyeGPT: Ophthalmic Assistant with Large Language Models
EyeGPT is a specialized large language model for ophthalmology that integrates role-playing, fine-tuning, and retrieval-augmented generation to enhance clinical performance. It achieves human-level understandability, trustworthiness, and empathy in ophthalmic consultations, with hallucination rates significantly reduced compared to general LLMs.
Artificial intelligence (AI) has gained significant attention in healthcare consultation due to its potential to improve clinical workflow and enhance medical communication. However, owing to the complex nature of medical information, large language models (LLM) trained with general world knowledge might not possess the capability to tackle medical-related tasks at an expert level. Here, we introduce EyeGPT, a specialized LLM designed specifically for ophthalmology, using three optimization strategies including role-playing, finetuning, and retrieval-augmented generation. In particular, we proposed a comprehensive evaluation framework that encompasses a diverse dataset, covering various subspecialties of ophthalmology, different users, and diverse inquiry intents. Moreover, we considered multiple evaluation metrics, including accuracy, understandability, trustworthiness, empathy, and the proportion of hallucinations. By assessing the performance of different EyeGPT variants, we identify the most effective one, which exhibits comparable levels of understandability, trustworthiness, and empathy to human ophthalmologists (all Ps>0.05). Overall, ur study provides valuable insights for future research, facilitating comprehensive comparisons and evaluations of different strategies for developing specialized LLMs in ophthalmology. The potential benefits include enhancing the patient experience in eye care and optimizing ophthalmologists' services.
Motivation & Objective
- To develop a specialized large language model tailored for ophthalmology that outperforms general-purpose LLMs in clinical relevance and accuracy.
- To address the limitations of general LLMs in handling complex, domain-specific medical information in ophthalmology.
- To design and evaluate a comprehensive framework for assessing ophthalmic LLMs across multiple dimensions, including accuracy, empathy, and hallucination rates.
- To identify the optimal combination of optimization strategies—role-playing, fine-tuning, and retrieval-augmented generation—for clinical deployment.
- To provide a benchmark for future research in specialized medical LLMs by introducing a diverse, multi-intent evaluation dataset across ophthalmology subspecialties.
Proposed method
- Employed role-playing prompting to simulate expert ophthalmologist behavior during inference.
- Applied domain-specific fine-tuning on a curated dataset of ophthalmic clinical queries and responses.
- Integrated retrieval-augmented generation (RAG) using a specialized ophthalmology knowledge base to improve factual consistency.
- Designed a multi-dimensional evaluation framework incorporating accuracy, understandability, trustworthiness, empathy, and hallucination detection.
- Utilized a diverse dataset spanning multiple ophthalmology subspecialties, user types, and inquiry intents to ensure robust evaluation.
- Evaluated model variants using both automatic metrics and human annotation to validate performance against human ophthalmologists.
Experimental results
Research questions
- RQ1What is the optimal combination of optimization strategies—role-playing, fine-tuning, and retrieval-augmented generation—for building a high-performing ophthalmic LLM?
- RQ2To what extent can a specialized LLM match human ophthalmologists in terms of understandability, trustworthiness, and empathy?
- RQ3How do different LLM variants perform in reducing hallucinations while maintaining clinical accuracy in ophthalmic consultations?
- RQ4Can a fine-tuned, retrieval-augmented LLM achieve comparable performance to human experts across diverse patient inquiry types and subspecialties?
- RQ5What are the key factors influencing the reliability and clinical usability of LLMs in ophthalmology?
Key findings
- The optimized EyeGPT variant achieved statistically indistinguishable levels of understandability, trustworthiness, and empathy compared to human ophthalmologists (all p > 0.05).
- The combination of role-playing, fine-tuning, and retrieval-augmented generation significantly reduced hallucination rates compared to baseline LLMs.
- EyeGPT demonstrated high accuracy across diverse ophthalmology subspecialties, including glaucoma, retinal disorders, and refractive surgery.
- Human evaluation confirmed that the model’s responses were rated as highly trustworthy and empathetic, with no significant difference from human expert responses.
- The comprehensive evaluation framework successfully captured nuanced performance dimensions beyond accuracy, including patient-centered communication qualities.
- The study provides a validated benchmark and methodology for developing and evaluating specialized LLMs in ophthalmology and other medical domains.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.