Skip to main content
QUICK REVIEW

[論文レビュー] Benchmarking ChatGPT-4 on ACR Radiation Oncology In-Training (TXIT) Exam and Red Journal Gray Zone Cases: Potentials and Challenges for AI-Assisted Medical Education and Decision Making in Radiation Oncology

Yixing Huang, Ahmed M. Gomaa|arXiv (Cornell University)|Apr 24, 2023
Artificial Intelligence in Healthcare and Education被引用数 9
ひとこと要約

この論文は、ACR TXIT 第38回試験および2022年 Red Journal Gray Zone ケースにおけるChatGPT-4をベンチマークし(ChatGPT-3.5との比較あり)、放射線腫瘍学の教育と意思決定におけるAIの可能性と課題を評価する。

ABSTRACT

The potential of large language models in medicine for education and decision making purposes has been demonstrated as they achieve decent scores on medical exams such as the United States Medical Licensing Exam (USMLE) and the MedQA exam. In this work, we evaluate the performance of ChatGPT-4 in the specialized field of radiation oncology using the 38th American College of Radiology (ACR) radiation oncology in-training (TXIT) exam and the 2022 Red Journal Gray Zone cases. For the TXIT exam, ChatGPT-3.5 and ChatGPT-4 have achieved the scores of 63.65% and 74.57%, respectively, highlighting the advantage of the latest ChatGPT-4 model. Based on the TXIT exam, ChatGPT-4's strong and weak areas in radiation oncology are identified to some extent. Specifically, ChatGPT-4 demonstrates better knowledge of statistics, CNS & eye, pediatrics, biology, and physics than knowledge of bone & soft tissue and gynecology, as per the ACR knowledge domain. Regarding clinical care paths, ChatGPT-4 performs better in diagnosis, prognosis, and toxicity than brachytherapy and dosimetry. It lacks proficiency in in-depth details of clinical trials. For the Gray Zone cases, ChatGPT-4 is able to suggest a personalized treatment approach to each case with high correctness and comprehensiveness. Importantly, it provides novel treatment aspects for many cases, which are not suggested by any human experts. Both evaluations demonstrate the potential of ChatGPT-4 in medical education for the general public and cancer patients, as well as the potential to aid clinical decision-making, while acknowledging its limitations in certain domains. Because of the risk of hallucination, facts provided by ChatGPT always need to be verified.

研究の動機と目的

  • 標準化された放射線腫瘍学試験(ACR TXIT 第38回)でのChatGPT-4の性能を評価し、知識領域ごとの強みとギャップを特定する。
  • Gray Zone臨床ケースでのChatGPT-4の能力を評価し、AI支援意思決定と教育的価値を測る。
  • ChatGPT-4とChatGPT-3.5を比較して改善点と残る限界を明らかにする。
  • 医用AI出力における幻覚リスクと検証の必要性を探る。
  • 放射線腫瘍学の教育および臨床意思決定支援への潜在的応用について示唆を提供する。

提案手法

  • 293問のTXIT(画像のみの問題を除く)でChatGPT-3.5とChatGPT-4をベンチマークし、領域別正答率を報告する。
  • 専門家投票を用いた初期および修正版のAI推奨のベンチマークとして、2022年Red Journal Gray Zoneケース(15ケース)を用いる。
  • 放射線腫瘍学の専門医としてChatGPT-4に促し、他の専門家の推奨を要約させ、整合性と更新動作を評価する。
  • 臨床専門家により、ChatGPT-4の初回および修正版の出力を正確さ・網羅性・新規性(幻覚の追跡を含む)で評価してもらう。
  • 専門家の意見に触れさせた後の初期と更新推奨を比較して、文脈学習効果を分析する。

実験結果

リサーチクエスチョン

  • RQ1放射線腫瘍学の知識領域および臨床ケアカテゴリごとに、TXIT試験でのChatGPT-4の成績はどうなるか。
  • RQ2Gray Zoneケースにおける個別化・網羅的・臨床的に正当化された推奨を生成するChatGPT-4の能力はどれほどで、人間の専門家とどう比較されるか。
  • RQ3放射線腫瘍学のタスクで、精度・網羅性・信頼性の点でChatGPT-3.5よりChatGPT-4が改善を示すか。
  • RQ4放射線腫瘍学における医療教育と意思決定支援での主な制約と幻覚リスクは何か。
  • RQ5専門家の意見を用いた文脈内学習は、Gray ZoneケースでChatGPT-4の推奨を幻覚を招くことなく改善できるか。

主な発見

  • ChatGPT-4は、初回TXIT評価で74.06%対63.14%、API再評価で78.77%対62.05%と、ChatGPT-3.5より高いTXIT正解率を達成した。
  • ドメイン別の性能では、統計、CNS・眼科、消化管、物理でChatGPT-4が優れる一方、骨/軟部組織およびリンパ腫/白血病では他ドメインに比べて劣り、婦人科は両モデルにとって依然として難しい。
  • 臨床ケア経路では、診断・治療決定・治療計画・毒性評価で60%以上の正答率を超えるが、近接治療(ブレイバース)と線量計測は60%を下回る。
  • Gray Zoneケースでは、初期推奨は概ね正確かつ網羅的。専門家の意見に触れた後、正確さと網羅性が向上し、幻覚は減少(初期平均正確性3.5、更新4.0;初期平均網羅性3.1、更新3.7)。
  • ChatGPT-4は人間の専門家が提案しない新規の治療要素を提案することが多く、臨床試験の詳細や局所再発の記述で妄想を生むことがあるため、検証が不可欠。
  • ChatGPT-4のGray Zone推奨は特定の専門家と一致する傾向があり、複数専門家の入力から利益を得る。平均整合投票は人間専門家の合意に近く(約25–29%程度)。
  • 本研究は、放射線腫瘍学における教育および意思決定支援のAIの潜在能力を強調する一方、幻覚、ドメインギャップ(例:近接治療/線量計測/試験)、医用画像解釈の限界といった substantial challenges を指摘している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。