Skip to main content
QUICK REVIEW

[논문 리뷰] Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4

Ting Fang Tan, Kabilan Elangovan|arXiv (Cornell University)|2024. 02. 15.
Artificial Intelligence in Healthcare and Education인용 수 6
한 줄 요약

본 논문은 안과 질문에 대해 여러 LLM 챗봇을 미세조정하고 GPT-4 기반 채점과 임상의 순위 간의 일치를 평가하여 높은 동의도를 보이고 일부 모델에서 임상적 부정확성을 강조합니다.

ABSTRACT

Purpose: To assess the alignment of GPT-4-based evaluation to human clinician experts, for the evaluation of responses to ophthalmology-related patient queries generated by fine-tuned LLM chatbots. Methods: 400 ophthalmology questions and paired answers were created by ophthalmologists to represent commonly asked patient questions, divided into fine-tuning (368; 92%), and testing (40; 8%). We find-tuned 5 different LLMs, including LLAMA2-7b, LLAMA2-7b-Chat, LLAMA2-13b, and LLAMA2-13b-Chat. For the testing dataset, additional 8 glaucoma QnA pairs were included. 200 responses to the testing dataset were generated by 5 fine-tuned LLMs for evaluation. A customized clinical evaluation rubric was used to guide GPT-4 evaluation, grounded on clinical accuracy, relevance, patient safety, and ease of understanding. GPT-4 evaluation was then compared against ranking by 5 clinicians for clinical alignment. Results: Among all fine-tuned LLMs, GPT-3.5 scored the highest (87.1%), followed by LLAMA2-13b (80.9%), LLAMA2-13b-chat (75.5%), LLAMA2-7b-Chat (70%) and LLAMA2-7b (68.8%) based on the GPT-4 evaluation. GPT-4 evaluation demonstrated significant agreement with human clinician rankings, with Spearman and Kendall Tau correlation coefficients of 0.90 and 0.80 respectively; while correlation based on Cohen Kappa was more modest at 0.50. Notably, qualitative analysis and the glaucoma sub-analysis revealed clinical inaccuracies in the LLM-generated responses, which were appropriately identified by the GPT-4 evaluation. Conclusion: The notable clinical alignment of GPT-4 evaluation highlighted its potential to streamline the clinical evaluation of LLM chatbot responses to healthcare-related queries. By complementing the existing clinician-dependent manual grading, this efficient and automated evaluation could assist the validation of future developments in LLM applications for healthcare.

연구 동기 및 목표

  • 미세조정된 LLM 챗봇이 생성한 안과 Q&A에 대한 GPT-4 기반 평가와 인간 임상의 전문가 간의 일치 여부를 평가한다.
  • 안과 환자 질문에 대해 다수의 LLM을 미세조정하여 벤치마크 데이터세트를 생성한다.
  • 정확도, 관련성, 안전성, 이해 용이성을 포함하는 GPT-4 주도 임상 루브릭을 사용하여 응답을 평가한다.

제안 방법

  • 일반 환자 질의 대표하는 안과 의사가 작성한 400개의 안과 질문과 짝지어진 답변을 생성한다.
  • 미세조정(368개 질문) 및 테스트(40개 질문) 세트로 분할하고, 테스트에 8개 유 glaucoma Q&A 쌍을 포함한다.
  • LLAMA2-7b, LLAMA2-7b-Chat, LLAMA2-13b, LLAMA2-13b-Chat (and variations) 를 미세조정한다.
  • 다섯 개의 미세조정된 LLM으로부터 테스트 데이터세트에 대한 200개의 응답을 생성하여 평가한다.
  • 임상 정확도, 관련성, 환자 안전 및 이해 용이성을 포함하는 맞춤형 임상 평가 루브릭을 사용하여 GPT-4 평가를 안내한다.
  • GPT-4 평가 결과를 다섯 명의 임상의 순위와 비교하여 임상 정합성을 평가한다.

실험 결과

연구 질문

  • RQ1미세조정된 LLM이 생성한 안과 Q&A를 판단할 때 GPT-4 평가가 인간 임상의 순위와 얼마나 잘 일치하는가?
  • RQ2어떤 LLM 구성들이 안과 질문 응답에서 최고의 임상 정합성과 성능을 제공하는가?
  • RQ3LLM이 임상적으로 어디에서 부족한가(예: glaucoma) 및 GPT-4가 부정확성을 얼마나 효과적으로 식별하는가?

주요 결과

  • GPT-4 평가가 인간 임상가 순위와 높은 일치를 보였으며, Spearman 0.90 및 Kendall Tau 0.80이다.
  • Cohen Kappa 일치도는 GPT-4 평가와 임상의 순위 사이에서 0.50으로 더 보수적이었다.
  • GPT-4 평가가 LLM 생성 응답에서 임상적 부정확성을 식별했으며, glaucoma 관련 문제를 포함한다.
  • 모든 미세조정 모델 중에서 GPT-3.5가 GPT-4 평가에서 87.1%로 최고점을 받았고, 그다음 LLAMA2-13b 80.9%, LLAMA2-13b-chat 75.5%, LLAMA2-7b-Chat 70%, LLAMA2-7b 68.8%였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.