Skip to main content
QUICK REVIEW

[Paper Review] ChatGPT Makes Medicine Easy to Swallow: An Exploratory Case Study on Simplified Radiology Reports

Katharina Jeblick, Balthasar Schachtner|arXiv (Cornell University)|Dec 30, 2022
Artificial Intelligence in Healthcare and Education105 citations
TL;DR

The study assesses radiologists’ quality judgments of ChatGPT-simplified radiology reports, finding they are largely factually correct and complete but show instances of inaccuracies and potentially harmful implications.

ABSTRACT

The release of ChatGPT, a language model capable of generating text that appears human-like and authentic, has gained significant attention beyond the research community. We expect that the convincing performance of ChatGPT incentivizes users to apply it to a variety of downstream tasks, including prompting the model to simplify their own medical reports. To investigate this phenomenon, we conducted an exploratory case study. In a questionnaire, we asked 15 radiologists to assess the quality of radiology reports simplified by ChatGPT. Most radiologists agreed that the simplified reports were factually correct, complete, and not potentially harmful to the patient. Nevertheless, instances of incorrect statements, missed key medical findings, and potentially harmful passages were reported. While further studies are needed, the initial insights of this study indicate a great potential in using large language models like ChatGPT to improve patient-centered care in radiology and other medical domains.

Motivation & Objective

  • Assess whether ChatGPT-simplified radiology reports are factually correct, complete, and safe for patients.
  • Investigate common error types and potential harm arising from automated simplifications.
  • Provide initial insights into opportunities and challenges of using LLMs for patient-centered radiology communication.

Proposed method

  • Design an exploratory case study with three fictitious radiology reports written by an experienced radiologist.
  • Prompt ChatGPT to create 15 unique simplified versions for each original report, totaling 45 outputs.
  • Have 15 radiologists rate the simplified reports on factual correctness, completeness, and potential harm using a structured questionnaire.
  • Analyze ratings with descriptive statistics (median, quantiles, IQR, min/max, mean, SD) and perform inductive free-text categorization of responses.

Experimental results

Research questions

  • RQ1What is radiologists’ opinion on the quality of ChatGPT-generated simplified radiology reports?
  • RQ2Are the simplified reports factually correct and complete, and do they pose any potential harm to patients?
  • RQ3What are the common error types or omissions in ChatGPT-generated simplifications?
  • RQ4How do ratings differ across the three original report types ( Knee MRI, Head MRI, Oncol. CT )?

Key findings

  • Radiologists generally agreed that simplified reports are factually correct and complete (median = 2 for both).
  • Potential harm ratings showed more variability (median = 4) with neutral and Agree responses present.
  • Free-text analysis revealed errors such as misinterpretation of medical terms, imprecise language, hallucinations, and missing key findings in several simplified reports.
  • Misinterpretations included differential diagnoses being presented as final diagnoses and terms like thyroid struma being misdescribed.
  • Word counts of simplified Knee MRI reports tended to be longer than originals (median 414 vs 222).
  • Across reports, incorrect passages were highlighted by 51% of participants, missing information by 22%, and potentially harmful conclusions by 36%.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.