[论文解读] Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4
该论文对若干用于眼科问题的LLM聊天机器人进行微调,并评估基于GPT-4的评分与临床医生排名之间的对应关系,显示出高度一致性并在某些模型中突出显示临床上的不准确之处。
Purpose: To assess the alignment of GPT-4-based evaluation to human clinician experts, for the evaluation of responses to ophthalmology-related patient queries generated by fine-tuned LLM chatbots. Methods: 400 ophthalmology questions and paired answers were created by ophthalmologists to represent commonly asked patient questions, divided into fine-tuning (368; 92%), and testing (40; 8%). We find-tuned 5 different LLMs, including LLAMA2-7b, LLAMA2-7b-Chat, LLAMA2-13b, and LLAMA2-13b-Chat. For the testing dataset, additional 8 glaucoma QnA pairs were included. 200 responses to the testing dataset were generated by 5 fine-tuned LLMs for evaluation. A customized clinical evaluation rubric was used to guide GPT-4 evaluation, grounded on clinical accuracy, relevance, patient safety, and ease of understanding. GPT-4 evaluation was then compared against ranking by 5 clinicians for clinical alignment. Results: Among all fine-tuned LLMs, GPT-3.5 scored the highest (87.1%), followed by LLAMA2-13b (80.9%), LLAMA2-13b-chat (75.5%), LLAMA2-7b-Chat (70%) and LLAMA2-7b (68.8%) based on the GPT-4 evaluation. GPT-4 evaluation demonstrated significant agreement with human clinician rankings, with Spearman and Kendall Tau correlation coefficients of 0.90 and 0.80 respectively; while correlation based on Cohen Kappa was more modest at 0.50. Notably, qualitative analysis and the glaucoma sub-analysis revealed clinical inaccuracies in the LLM-generated responses, which were appropriately identified by the GPT-4 evaluation. Conclusion: The notable clinical alignment of GPT-4 evaluation highlighted its potential to streamline the clinical evaluation of LLM chatbot responses to healthcare-related queries. By complementing the existing clinician-dependent manual grading, this efficient and automated evaluation could assist the validation of future developments in LLM applications for healthcare.
研究动机与目标
- 评估由微调的LLM聊天机器人生成的眼科问答中,基于GPT-4的评估与人类临床专家之间的一致性。
- 对眼科患者问题对多个LLM进行微调,以创建一个基准数据集。
- 使用基于GPT-4的临床评估量表对回答进行评估,涵盖准确性、相关性、安全性和易懂性。
提出的方法
- 由眼科医生创建代表常见患者提问的400个眼科问题及配对答案。
- 分为微调集(368个问题)和测试集(40个问题);在测试集中包含8对青光眼的问答。
- 对五个LLM进行微调:LLAMA2-7b、LLAMA2-7b-Chat、LLAMA2-13b、LLAMA2-13b-Chat(及其变体)。
- 从五个微调后的LLM生成测试数据集的200个回答以供评估。
- 使用定制的临床评估量表来引导GPT-4评估,重点关注临床准确性、相关性、患者安全性和易理解性。
- 将GPT-4评估结果与五位临床医生的排名进行比较,以评估临床一致性。
实验结果
研究问题
- RQ1GPT-4评估在判断由微调的LLM生成的眼科问答时,与人类临床医生的排名的一致性有多高?
- RQ2哪些LLM配置在眼科问答中提供最佳的临床一致性和性能?
- RQ3LLMs在临床上哪里存在不足(如青光眼),GPT-4发现不准确之处的效果如何?
主要发现
- GPT-4评估与人类临床医生排名的一致性很高,Spearman 0.90,Kendall Tau 0.80。
- Cohen Kappa一致性在GPT-4评估与临床医生排名之间为0.50,较为温和。
- GPT-4评估发现了LLM生成回答中的临床不准确性,包括与青光眼相关的问题。
- 在所有微调模型中,GPT-3.5在GPT-4评估中的得分最高,为87.1%,其次是LLAMA2-13b 80.9%,LLAMA2-13b-chat 75.5%,LLAMA2-7b-Chat 70%,LLAMA2-7b 68.8%。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。