Skip to main content
QUICK REVIEW

[Paper Review] Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4

Ting Fang Tan, Kabilan Elangovan|arXiv (Cornell University)|Feb 15, 2024
Artificial Intelligence in Healthcare and Education6 citations
TL;DR

The paper fine-tunes several LLM chatbots for ophthalmology questions and evaluates GPT-4-based scoring against clinician rankings, showing high agreement and highlighting clinical inaccuracies in some models.

ABSTRACT

Purpose: To assess the alignment of GPT-4-based evaluation to human clinician experts, for the evaluation of responses to ophthalmology-related patient queries generated by fine-tuned LLM chatbots. Methods: 400 ophthalmology questions and paired answers were created by ophthalmologists to represent commonly asked patient questions, divided into fine-tuning (368; 92%), and testing (40; 8%). We find-tuned 5 different LLMs, including LLAMA2-7b, LLAMA2-7b-Chat, LLAMA2-13b, and LLAMA2-13b-Chat. For the testing dataset, additional 8 glaucoma QnA pairs were included. 200 responses to the testing dataset were generated by 5 fine-tuned LLMs for evaluation. A customized clinical evaluation rubric was used to guide GPT-4 evaluation, grounded on clinical accuracy, relevance, patient safety, and ease of understanding. GPT-4 evaluation was then compared against ranking by 5 clinicians for clinical alignment. Results: Among all fine-tuned LLMs, GPT-3.5 scored the highest (87.1%), followed by LLAMA2-13b (80.9%), LLAMA2-13b-chat (75.5%), LLAMA2-7b-Chat (70%) and LLAMA2-7b (68.8%) based on the GPT-4 evaluation. GPT-4 evaluation demonstrated significant agreement with human clinician rankings, with Spearman and Kendall Tau correlation coefficients of 0.90 and 0.80 respectively; while correlation based on Cohen Kappa was more modest at 0.50. Notably, qualitative analysis and the glaucoma sub-analysis revealed clinical inaccuracies in the LLM-generated responses, which were appropriately identified by the GPT-4 evaluation. Conclusion: The notable clinical alignment of GPT-4 evaluation highlighted its potential to streamline the clinical evaluation of LLM chatbot responses to healthcare-related queries. By complementing the existing clinician-dependent manual grading, this efficient and automated evaluation could assist the validation of future developments in LLM applications for healthcare.

Motivation & Objective

  • Assess alignment between GPT-4-based evaluation and human clinician experts for ophthalmology Q&A generated by fine-tuned LLM chatbots.
  • Fine-tune multiple LLMs on ophthalmology patient questions to create a benchmark dataset.
  • Evaluate responses using a GPT-4-driven clinical rubric across accuracy, relevance, safety, and understandability.

Proposed method

  • Create 400 ophthalmology questions and paired answers by ophthalmologists representing common patient queries.
  • Split into fine-tuning (368 questions) and testing (40 questions) sets; include 8 glaucoma Q&A pairs in testing.
  • Fine-tune five LLMs: LLAMA2-7b, LLAMA2-7b-Chat, LLAMA2-13b, LLAMA2-13b-Chat (and variations).
  • Generate 200 responses to the testing dataset from the five fine-tuned LLMs for evaluation.
  • Use a customized clinical evaluation rubric to guide GPT-4 evaluation focusing on clinical accuracy, relevance, patient safety, and ease of understanding.
  • Compare GPT-4 evaluation results with rankings by five clinicians to assess clinical alignment.

Experimental results

Research questions

  • RQ1How well does GPT-4 evaluation align with human clinician rankings in judging ophthalmology Q&A from fine-tuned LLMs?
  • RQ2Which LLM configurations provide the best clinical alignment and performance in ophthalmology question answering?
  • RQ3Where do LLMs fall short clinically (e.g., glaucoma) and how effectively does GPT-4 identify inaccuracies?

Key findings

  • GPT-4 evaluation achieved high agreement with human clinician rankings, with Spearman 0.90 and Kendall Tau 0.80.
  • Cohen Kappa agreement between GPT-4 evaluation and clinician rankings was more modest at 0.50.
  • GPT-4 evaluation identified clinical inaccuracies in LLM-generated responses, including glaucoma-specific issues.
  • Among all fine-tuned models, GPT-3.5 scored the highest at 87.1% in GPT-4 evaluation, followed by LLAMA2-13b at 80.9%, LLAMA2-13b-chat at 75.5%, LLAMA2-7b-Chat at 70%, and LLAMA2-7b at 68.8%.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.