Skip to main content
QUICK REVIEW

[论文解读] Benchmarking ChatGPT-4 on ACR Radiation Oncology In-Training (TXIT) Exam and Red Journal Gray Zone Cases: Potentials and Challenges for AI-Assisted Medical Education and Decision Making in Radiation Oncology

Yixing Huang, Ahmed M. Gomaa|arXiv (Cornell University)|Apr 24, 2023
Artificial Intelligence in Healthcare and Education被引用 9
一句话总结

这篇论文在第38届ACR TXIT考试和2022年Red Journal Gray Zone病例上对ChatGPT-4进行了基准测试(并与ChatGPT-3.5进行比较),以评估AI在放射肿瘤教育与决策中的潜力与挑战。

ABSTRACT

The potential of large language models in medicine for education and decision making purposes has been demonstrated as they achieve decent scores on medical exams such as the United States Medical Licensing Exam (USMLE) and the MedQA exam. In this work, we evaluate the performance of ChatGPT-4 in the specialized field of radiation oncology using the 38th American College of Radiology (ACR) radiation oncology in-training (TXIT) exam and the 2022 Red Journal Gray Zone cases. For the TXIT exam, ChatGPT-3.5 and ChatGPT-4 have achieved the scores of 63.65% and 74.57%, respectively, highlighting the advantage of the latest ChatGPT-4 model. Based on the TXIT exam, ChatGPT-4's strong and weak areas in radiation oncology are identified to some extent. Specifically, ChatGPT-4 demonstrates better knowledge of statistics, CNS & eye, pediatrics, biology, and physics than knowledge of bone & soft tissue and gynecology, as per the ACR knowledge domain. Regarding clinical care paths, ChatGPT-4 performs better in diagnosis, prognosis, and toxicity than brachytherapy and dosimetry. It lacks proficiency in in-depth details of clinical trials. For the Gray Zone cases, ChatGPT-4 is able to suggest a personalized treatment approach to each case with high correctness and comprehensiveness. Importantly, it provides novel treatment aspects for many cases, which are not suggested by any human experts. Both evaluations demonstrate the potential of ChatGPT-4 in medical education for the general public and cancer patients, as well as the potential to aid clinical decision-making, while acknowledging its limitations in certain domains. Because of the risk of hallucination, facts provided by ChatGPT always need to be verified.

研究动机与目标

  • 评估ChatGPT-4在标准化放射肿瘤学考试(ACR TXIT 第38届)上的表现,并在知识领域内识别优势与差距。
  • 评估ChatGPT-4在Gray Zone临床病例上的能力,以衡量AI辅助决策和教育价值。
  • 将ChatGPT-4与ChatGPT-3.5进行比较,以确定改进及仍然存在的局限性。
  • 探索AI医疗输出中的幻觉风险以及需要验证的问题。
  • 提供放射肿瘤学在教育和临床决策支持方面的潜在应用洞察。

提出的方法

  • 对293道TXIT题(不含仅图片题)基准测试ChatGPT-3.5和ChatGPT-4,并报道各知识领域的准确率。
  • 使用2022年Red Journal Gray Zone病例(15例)并进行专家投票,作为初始与修订AI建议的基准。
  • 将ChatGPT-4设定为放射肿瘤科专家,概述其他专家的建议,并评估对齐性及更新行为。
  • 请临床专家对ChatGPT-4的初始与修订输出在正确性、全面性及新颖性(同时跟踪幻觉)进行评分。
  • 通过暴露于专家意见后比较初始与更新的建议,分析上下文学习效应。

实验结果

研究问题

  • RQ1ChatGPT-4在TXIT考试中在不同放射肿瘤学知识领域和临床护理类别的表现如何?
  • RQ2ChatGPT-4在Gray Zone病例中生成个性化、全面且临床上可辩护的建议的能力如何,与人类专家相比如何?
  • RQ3ChatGPT-4在放射肿瘤学任务的准确性、全面性和可靠性方面是否相较于ChatGPT-3.5有所提升?
  • RQ4在放射肿瘼学的AI辅助医学教育与决策中,主要的局限性与幻觉风险有哪些?
  • RQ5在Gray Zone病例中,是否可以通过引入专家意见的上下文学习来提升ChatGPT-4的建议,同时不引入幻觉?

主要发现

  • ChatGPT-4在TXIT初步评估中的准确率(74.06%)高于ChatGPT-3.5(63.14%);在API重复评估中的准确率为78.77%对比62.05%。
  • 领域表现显示ChatGPT-4在统计、CNS与眼科、消化系统以及物理领域的表现优于某些领域;但相对于某些领域,在骨/软组织和淋巴瘤/白血病方面表现不足;妇科对两种模型仍具挑战。
  • 在临床护理路径中,ChatGPT-4在诊断、治疗决策、治疗计划和毒性方面的准确率超过60%,但在近距离放射治疗(brachytherapy)和剂量学上存在困难(低于60%)。
  • 在Gray Zone病例中,ChatGPT-4的初始建议通常正确且全面;在暴露于专家意见后,准确性和全面性提升,幻觉减少(初始平均正确性3.5,更新4.0;初始平均全面性3.1,更新3.7)。
  • ChatGPT-4经常提出人类专家未建议的治疗新点,且在临床试验细节或局部复发陈述方面可能产生幻觉;需要进行核实。
  • ChatGPT-4在Gray Zone的建议往往与部分专家保持一致,并且多专家输入带来好处,平均对齐投票接近人类专家共识(大约25–29%)。
  • 研究强调AI在放射肿瘤学中的教育与决策支持潜力,同时面临包括幻觉、领域差距(如近距离放疗/剂量学/试验)以及医学影像解读的局限性等重大挑战。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。