Skip to main content
QUICK REVIEW

[Paper Review] Benchmarking ChatGPT-4 on ACR Radiation Oncology In-Training (TXIT) Exam and Red Journal Gray Zone Cases: Potentials and Challenges for AI-Assisted Medical Education and Decision Making in Radiation Oncology

Yixing Huang, Ahmed M. Gomaa|arXiv (Cornell University)|Apr 24, 2023
Artificial Intelligence in Healthcare and Education9 citations
TL;DR

This paper benchmarks ChatGPT-4 (and compares to ChatGPT-3.5) on the 38th ACR TXIT exam and 2022 Red Journal Gray Zone cases to assess potential and challenges of AI for radiation oncology education and decision-making.

ABSTRACT

The potential of large language models in medicine for education and decision making purposes has been demonstrated as they achieve decent scores on medical exams such as the United States Medical Licensing Exam (USMLE) and the MedQA exam. In this work, we evaluate the performance of ChatGPT-4 in the specialized field of radiation oncology using the 38th American College of Radiology (ACR) radiation oncology in-training (TXIT) exam and the 2022 Red Journal Gray Zone cases. For the TXIT exam, ChatGPT-3.5 and ChatGPT-4 have achieved the scores of 63.65% and 74.57%, respectively, highlighting the advantage of the latest ChatGPT-4 model. Based on the TXIT exam, ChatGPT-4's strong and weak areas in radiation oncology are identified to some extent. Specifically, ChatGPT-4 demonstrates better knowledge of statistics, CNS & eye, pediatrics, biology, and physics than knowledge of bone & soft tissue and gynecology, as per the ACR knowledge domain. Regarding clinical care paths, ChatGPT-4 performs better in diagnosis, prognosis, and toxicity than brachytherapy and dosimetry. It lacks proficiency in in-depth details of clinical trials. For the Gray Zone cases, ChatGPT-4 is able to suggest a personalized treatment approach to each case with high correctness and comprehensiveness. Importantly, it provides novel treatment aspects for many cases, which are not suggested by any human experts. Both evaluations demonstrate the potential of ChatGPT-4 in medical education for the general public and cancer patients, as well as the potential to aid clinical decision-making, while acknowledging its limitations in certain domains. Because of the risk of hallucination, facts provided by ChatGPT always need to be verified.

Motivation & Objective

  • Assess ChatGPT-4 performance on a standardized radiation oncology exam (ACR TXIT 38th) and identify strengths and gaps across knowledge domains.
  • Evaluate ChatGPT-4 capabilities on Gray Zone clinical cases to gauge AI-assisted decision making and educational value.
  • Compare ChatGPT-4 with ChatGPT-3.5 to determine improvements and remaining limitations.
  • Explore risks of hallucinations and the need for verification in medical AI outputs.
  • Provide insights into potential educational and clinical decision-support applications for radiation oncology.

Proposed method

  • Benchmark ChatGPT-3.5 and ChatGPT-4 on 293 TXIT questions (excluding image-only questions) and report domain-specific accuracies.
  • Use the 2022 Red Journal Gray Zone cases (15 cases) with expert voting as a benchmark for initial and revised AI recommendations.
  • Prompt ChatGPT-4 as an expert radiation oncologist, summarize other experts’ recommendations, and assess alignment and update behavior.
  • Have clinical experts rate ChatGPT-4’s initial and revised outputs on correctness, comprehensiveness, and novelty (hallucinations tracked).
  • Analyze in-context learning effects by comparing initial versus updated recommendations after exposure to expert opinions.

Experimental results

Research questions

  • RQ1How does ChatGPT-4 perform on the TXIT exam across different radiation oncology knowledge domains and clinical-care categories?
  • RQ2What is ChatGPT-4’s ability to generate personalized, comprehensive, and clinically justifiable recommendations for Gray Zone cases, and how does it compare to human experts?
  • RQ3Does ChatGPT-4 exhibit improvements over ChatGPT-3.5 in accuracy, comprehensiveness, and reliability on radiation oncology tasks?
  • RQ4What are the main limitations and hallucination risks in AI-assisted medical education and decision making within radiation oncology?
  • RQ5Can in-context learning with expert opinions improve ChatGPT-4’s recommendations in Gray Zone cases without introducing hallucinations?

Key findings

  • ChatGPT-4 achieves higher TXIT accuracy than ChatGPT-3.5 (74.06% vs 63.14% in initial TXIT assessment; 78.77% vs 62.05% in API-repeated assessment).
  • Domain performance shows ChatGPT-4 outperforms in statistics, CNS & eye, GI, and physics, but underperforms in bone/soft tissue and lymphoma/leukemia relative to some domains; gynecology remains challenging for both models.
  • In clinical-care paths, ChatGPT-4 exceeds 60% accuracy for diagnosis, treatment decision, treatment planning, and toxicity, but struggles with brachytherapy and dosimetry (below 60%).
  • On Gray Zone cases, ChatGPT-4’s initial recommendations are generally correct and comprehensive; after exposure to expert opinions, accuracy and comprehensiveness improve, with fewer hallucinations (initial mean correctness 3.5, updated 4.0; initial mean comprehensiveness 3.1, updated 3.7).
  • ChatGPT-4 often proposes novel treatment aspects not suggested by human experts and can hallucinate in clinical trial details or local recurrence statements; verification is essential.
  • ChatGPT-4’s Gray Zone recommendations tend to align with certain experts and benefit from multi-expert input, with average alignment votes near human expert consensus (around 25–29%).
  • The study highlights educational and decision-support potential of AI in radiation oncology, alongside substantial challenges including hallucinations, domain gaps (e.g., brachytherapy/dosimetry/trials), and limitations in medical image interpretation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.