Skip to main content
QUICK REVIEW

[论文解读] Evaluating multiple large language models in pediatric ophthalmology

Jason Holmes, Rui Peng|arXiv (Cornell University)|Nov 7, 2023
Artificial Intelligence in Healthcare and Education被引用 4
一句话总结

本研究通过100道多项选择题的考试,评估了三种大型语言模型——GPT-4、ChatGPT(GPT-3.5)和PaLM2——在儿科眼科领域与医学生、研究生及主治医师的表现。GPT-4在准确率上与主治医师相当,优于医学生,并在响应稳定性与置信度方面优于其他模型。

ABSTRACT

IMPORTANCE The response effectiveness of different large language models (LLMs) and various individuals, including medical students, graduate students, and practicing physicians, in pediatric ophthalmology consultations, has not been clearly established yet. OBJECTIVE Design a 100-question exam based on pediatric ophthalmology to evaluate the performance of LLMs in highly specialized scenarios and compare them with the performance of medical students and physicians at different levels. DESIGN, SETTING, AND PARTICIPANTS This survey study assessed three LLMs, namely ChatGPT (GPT-3.5), GPT-4, and PaLM2, were assessed alongside three human cohorts: medical students, postgraduate students, and attending physicians, in their ability to answer questions related to pediatric ophthalmology. It was conducted by administering questionnaires in the form of test papers through the LLM network interface, with the valuable participation of volunteers. MAIN OUTCOMES AND MEASURES Mean scores of LLM and humans on 100 multiple-choice questions, as well as the answer stability, correlation, and response confidence of each LLM. RESULTS GPT-4 performed comparably to attending physicians, while ChatGPT (GPT-3.5) and PaLM2 outperformed medical students but slightly trailed behind postgraduate students. Furthermore, GPT-4 exhibited greater stability and confidence when responding to inquiries compared to ChatGPT (GPT-3.5) and PaLM2. CONCLUSIONS AND RELEVANCE Our results underscore the potential for LLMs to provide medical assistance in pediatric ophthalmology and suggest significant capacity to guide the education of medical students.

研究动机与目标

  • 评估大型语言模型(LLMs)在儿科眼科领域临床推理准确性的目标。
  • 将大型语言模型的表现与医学生、研究生及主治医师进行比较。
  • 评估大型语言模型与人类群体在响应稳定性、置信度及相关性方面的表现。
  • 探讨大型语言模型在专业医学领域中的教育与临床辅助潜力。

提出的方法

  • 基于儿科眼科内容开发了一套包含100道题的多项选择题考试。
  • 通过其网络接口查询了三种大型语言模型——ChatGPT(GPT-3.5)、GPT-4和PaLM2——以回答考试题目。
  • 三个由医学生、研究生及主治医师组成的群体完成了相同的考试。
  • 通过平均得分、响应稳定性、置信度水平以及与人类回答的相关性来衡量表现。
  • 通过统计分析比较了大型语言模型与人类群体在关键指标上的表现。

实验结果

研究问题

  • RQ1不同大型语言模型在儿科眼科领域与临床医生相比表现如何?
  • RQ2大型语言模型能否达到甚至超过医学生和医生的诊断准确性?
  • RQ3与人类响应相比,大型语言模型的响应在稳定性和置信度方面如何?
  • RQ4大型语言模型与人类在临床推理任务中的表现相关性如何?

主要发现

  • GPT-4的平均得分与主治医师相当,表明其在儿科眼科领域具有高水平的临床推理准确性。
  • ChatGPT(GPT-3.5)和PaLM2优于医学生,但得分略低于研究生群体。
  • GPT-4在响应稳定性方面优于ChatGPT(GPT-3.5)和PaLM2,并展现出更高的置信度得分。
  • GPT-4与主治医师之间的表现相关性强于其他大型语言模型与人类配对之间的相关性。
  • 大型语言模型,尤其是GPT-4,在100道题的测试集上表现出一致且可靠的表现。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。