[论文解读] Performance of ChatGPT on USMLE: Unlocking the Potential of Large Language Models for AI-Assisted Medical Education
本研究在使用哈佛解剖数据和医师裁定的前提下,评估了 ChatGPT 在 USMLE 风格题目上的表现,结果显示 ChatGPT 相较于 Google 更注重上下文关联和演绎推理,逻辑题正确率为 58.8%,伦理题为 60%。
Artificial intelligence is gaining traction in more ways than ever before. The popularity of language models and AI-based businesses has soared since ChatGPT was made available to the general public via OpenAI. It is becoming increasingly common for people to use ChatGPT both professionally and personally. Considering the widespread use of ChatGPT and the reliance people place on it, this study determined how reliable ChatGPT can be for answering complex medical and clinical questions. Harvard University gross anatomy along with the United States Medical Licensing Examination (USMLE) questionnaire were used to accomplish the objective. The paper evaluated the obtained results using a 2-way ANOVA and posthoc analysis. Both showed systematic covariation between format and prompt. Furthermore, the physician adjudicators independently rated the outcome's accuracy, concordance, and insight. As a result of the analysis, ChatGPT-generated answers were found to be more context-oriented and represented a better model for deductive reasoning than regular Google search results. Furthermore, ChatGPT obtained 58.8% on logical questions and 60% on ethical questions. This means that the ChatGPT is approaching the passing range for logical questions and has crossed the threshold for ethical questions. The paper believes ChatGPT and other language learning models can be invaluable tools for e-learners; however, the study suggests that there is still room to improve their accuracy. In order to improve ChatGPT's performance in the future, further research is needed to better understand how it can answer different types of questions.
研究动机与目标
- 评估 ChatGPT 对回答与 USMLE‑风格评估相关的复杂医学与临床问题的可靠性。
- 将 ChatGPT 的表现与传统搜索方法(Google)在医疗问答格式上的表现进行比较。
- 使用统计分析评估题目格式与提示对 ChatGPT 表现的影响。
- 让医师裁判独立评估 AI 生成答案的准确性、一致性和洞察力。
提出的方法
- 以哈佛大学解剖学内容和 USMLE‑风格题目作为评估材料。
- 应用两因素方差分析(2‑WAY ANOVA)来分析格式与提示对表现的影响。
- 进行事后分析以探讨格式与提示之间的交互效应。
- 让医师裁判独立对 AI 输出在准确性、一致性和洞察力方面进行评分。
实验结果
研究问题
- RQ1ChatGPT 是否能够在逻辑和伦理领域可靠地回答 USMLE 风格的问题?
- RQ2题目格式或提示风格是否系统性地影响 ChatGPT 的表现?
- RQ3医师评估者如何相较于标准搜索结果对 ChatGPT 回答的准确性、一致性和洞察力进行评分?
主要发现
- ChatGPT 生成的答案比 Google 搜索结果更具上下文导向性。
- ChatGPT 在回答中表现出比常规 Google 搜索结果更强的演绎推理能力。
- ChatGPT 在逻辑题上的得分为 58.8%,在伦理题上为 60%。
- ChatGPT 似乎在逻辑题上接近及格区间,在伦理题上已超过阈值。
- 统计分析表明格式与提示之间存在系统性相关性,见两因素方差分析和事后分析。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。