[论文解读] Evaluating Large Language Models in Ophthalmology
本研究通过向大型语言模型和三组人类专业人员(本科生、硕士生、主治医师)发放100道多项选择题的考题,评估了GPT-3.5、GPT-4和PaLM2在眼科领域的表现。GPT-4在准确率、稳定性和置信度方面均优于所有人类群体,其表现与主治医师相当,显示出在医学教育和临床决策支持方面具有巨大潜力。
Purpose: The performance of three different large language models (LLMS) (GPT-3.5, GPT-4, and PaLM2) in answering ophthalmology professional questions was evaluated and compared with that of three different professional populations (medical undergraduates, medical masters, and attending physicians). Methods: A 100-item ophthalmology single-choice test was administered to three different LLMs (GPT-3.5, GPT-4, and PaLM2) and three different professional levels (medical undergraduates, medical masters, and attending physicians), respectively. The performance of LLM was comprehensively evaluated and compared with the human group in terms of average score, stability, and confidence. Results: Each LLM outperformed undergraduates in general, with GPT-3.5 and PaLM2 being slightly below the master's level, while GPT-4 showed a level comparable to that of attending physicians. In addition, GPT-4 showed significantly higher answer stability and confidence than GPT-3.5 and PaLM2. Conclusion: Our study shows that LLM represented by GPT-4 performs better in the field of ophthalmology. With further improvements, LLM will bring unexpected benefits in medical education and clinical decision making in the near future.
研究动机与目标
- 评估大型语言模型(LLMs)在眼科领域相对于人类医学专业人员的表现。
- 比较GPT-3.5、GPT-4和PaLM2在标准化眼科问题上的准确率、稳定性和置信度。
- 评估大型语言模型是否能够达到甚至超越医学生和主治医师的诊断推理能力。
- 探讨大型语言模型在医学教育和临床决策支持应用中的潜力。
提出的方法
- 开发了一套包含100道题的多项选择题眼科测试,用于评估核心眼科领域中的临床知识。
- 通过基于提示的推理方式,向三组大型语言模型(GPT-3.5、GPT-4和PaLM2)发放测试。
- 在受控条件下,由三组人类群体(本科生、硕士生、主治医师)完成相同的测试。
- 通过平均得分、答案稳定性(多次重复测试的一致性)以及自我置信度评分来衡量表现。
- 对大型语言模型与人类群体的表现进行统计比较,以评估相对表现。
- 收集每项大型语言模型预测的置信度评分,以评估其校准程度和自我评估的可靠性。
实验结果
研究问题
- RQ1GPT-3.5、GPT-4和PaLM2在标准化眼科知识测试中的表现与人类专业人员相比如何?
- RQ2GPT-4在眼科领域的表现是否与主治医师相当?
- RQ3大型语言模型在答案稳定性与自我置信度方面与人类群体相比如何?
- RQ4大型语言模型能否在准确率和一致性方面超越不同培训水平的医学生?
主要发现
- GPT-4的平均得分与主治医师相当,优于本科生和硕士生。
- GPT-3.5和PaLM2的得分略低于硕士生水平,表明其表现中等但未占优势。
- GPT-4的答案稳定性显著高于GPT-3.5和PaLM2,表明其推理更具一致性。
- GPT-4的自我置信度评分高于GPT-3.5和PaLM2,且与实际表现的校准更优。
- 所有大型语言模型在整体准确率上均优于本科生,凸显其在基础医学教育中的潜力。
- 结果表明,GPT-4在眼科知识方面与经验丰富的临床医生相当,具备出色的可靠性与置信度指标。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。