Skip to main content
QUICK REVIEW

[论文解读] GPT-4 to GPT-3.5: 'Hold My Scalpel' -- A Look at the Competency of OpenAI's GPT on the Plastic Surgery In-Service Training Exam

Jonathan D. Freedman, Ian A. Nappier|arXiv (Cornell University)|Apr 4, 2023
Artificial Intelligence in Healthcare and Education被引用 7
一句话总结

本研究评估了 GPT-4 和 GPT-3.5 在高风险临床基准测试——整形外科学住院医师培训考试(PSITE)上的表现。通过使用具有真实临床情景的原始多选题,作者证明 GPT-4 在 2022 年和 2021 年考试中分别取得了第 88 百分位数和第 99 百分位数的成绩,显著优于 GPT-3.5,后者分别仅取得第 8 百分位数和第 3 百分位数的成绩,表明其在医学推理能力方面实现了质的飞跃。

ABSTRACT

The Plastic Surgery In-Service Training Exam (PSITE) is an important indicator of resident proficiency and serves as a useful benchmark for evaluating OpenAI's GPT. Unlike many of the simulated tests or practice questions shown in the GPT-4 Technical Paper, the multiple-choice questions evaluated here are authentic PSITE questions. These questions offer realistic clinical vignettes that a plastic surgeon commonly encounters in practice and scores highly correlate with passing the written boards required to become a Board Certified Plastic Surgeon. Our evaluation shows dramatic improvement of GPT-4 (without vision) over GPT-3.5 with both the 2022 and 2021 exams respectively increasing the score from 8th to 88th percentile and 3rd to 99th percentile. The final results of the 2023 PSITE are set to be released on April 11, 2023, and this is an exciting moment to continue our research with a fresh exam. Our evaluation pipeline is ready for the moment that the exam is released so long as we have access via OpenAI to the GPT-4 API. With multimodal input, we may achieve superhuman performance on the 2023.

研究动机与目标

  • 评估 GPT-4 和 GPT-3.5 在真实世界、高风险医学考试中的临床推理能力。
  • 评估 GPT 模型是否能在整形外科学住院医师培训考试(PSITE)中达到与人类住院医师相当的水平。
  • 使用与执业认证成功相关的、真实的临床相关多选题,建立基准测试。
  • 研究 GPT-4 与 GPT-3.5 在复杂、专业领域医学考试中的表现差距。
  • 为 GPT-4 即将推出的 2023 年 PSITE 考试的未来评估做好准备。

提出的方法

  • 本研究使用 2022 年和 2021 年 PSITE 考试中的真实 PSITE 多选题作为评估数据集。
  • GPT-4 和 GPT-3.5 被每个问题提示,并需从提供的选项中选择最佳答案。
  • 性能通过与实际住院医师在相同考试中的表现相比的百分位数排名来衡量。
  • 评估流程设计为在 2023 年 PSITE 考试发布后可立即投入使用。
  • 本次评估未使用多模态输入,尽管作者指出未来若具备视觉能力,可能实现超越人类的表现。
  • 结果与历史住院医师表现的百分位数进行比较,以定位模型表现。

实验结果

研究问题

  • RQ1GPT-4 是否能在 PSITE 上达到或超过人类整形外科学住院医师的水平?
  • RQ2GPT-4 在同一临床推理考试中与 GPT-3.5 的表现相比如何?
  • RQ3GPT 模型在具有复杂临床情景的真实、高风险医学考试题目上的泛化能力如何?
  • RQ4GPT-4 与 GPT-3.5 之间的表现差距是否反映了医学推理能力的实质性提升,还是仅统计噪声?
  • RQ5未来对 GPT-4 在 2023 年 PSITE 考试中的评估,能否进一步揭示大语言模型在普通外科学领域的发展轨迹?

主要发现

  • GPT-4 在 2022 年 PSITE 考试中取得第 88 百分位数,表明其相对于人类住院医师的表现强劲。
  • GPT-4 在 2021 年 PSITE 考试中取得第 99 百分位数,表明其已达到接近专家水平的熟练度。
  • GPT-3.5 在 2022 年考试中仅取得第 8 百分位数,在 2021 年考试中仅取得第 3 百分位数,表现显著较低。
  • GPT-4 与 GPT-3.5 之间的表现差距显著,两者在两场考试中均超过 80 个百分点,GPT-4 显著领先。
  • 结果表明,GPT-4 已达到接近或超过整形外科学训练有素人类住院医师的临床推理能力水平。
  • 评估流程已准备就绪,可在 2023 年 PSITE 考试发布并可通过 OpenAI API 访问后立即用于评估 GPT-4。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。