[Paper Review] Translating Radiology Reports into Plain Language using ChatGPT and GPT-4 with Prompt Learning: Promising Results, Limitations, and Potential
The paper evaluates translating radiology reports to plain language using ChatGPT and GPT-4 with prompt learning, showing promising quality and useful suggestions but noting inconsistencies and remaining limitations.
The large language model called ChatGPT has drawn extensively attention because of its human-like expression and reasoning abilities. In this study, we investigate the feasibility of using ChatGPT in experiments on using ChatGPT to translate radiology reports into plain language for patients and healthcare providers so that they are educated for improved healthcare. Radiology reports from 62 low-dose chest CT lung cancer screening scans and 76 brain MRI metastases screening scans were collected in the first half of February for this study. According to the evaluation by radiologists, ChatGPT can successfully translate radiology reports into plain language with an average score of 4.27 in the five-point system with 0.08 places of information missing and 0.07 places of misinformation. In terms of the suggestions provided by ChatGPT, they are general relevant such as keeping following-up with doctors and closely monitoring any symptoms, and for about 37% of 138 cases in total ChatGPT offers specific suggestions based on findings in the report. ChatGPT also presents some randomness in its responses with occasionally over-simplified or neglected information, which can be mitigated using a more detailed prompt. Furthermore, ChatGPT results are compared with a newly released large model GPT-4, showing that GPT-4 can significantly improve the quality of translated reports. Our results show that it is feasible to utilize large language models in clinical education, and further efforts are needed to address limitations and maximize their potential.
Motivation & Objective
- Assess feasibility of translating radiology reports into plain language for patients and providers using ChatGPT and GPT-4.
- Evaluate the quality of translations and the usefulness of generated patient/provider suggestions.
- Investigate how prompt design affects translation quality and the role of prompt optimization and ensemble approaches.
Proposed method
- Collected 62 chest CT lung cancer screening reports and 76 brain MRI screening reports from a clinical database.
- Applied three prompts to ChatGPT: translation to plain language, patient suggestions, and provider suggestions.
- Compared ChatGPT translations with radiologist evaluations on completeness, correctness, and overall quality.
- Compared ChatGPT with GPT-4 using the same prompts and evaluation framework.
- Explored prompt optimization, prompt engineering variations, and ensemble translations to assess impact on quality.
Experimental results
Research questions
- RQ1Can ChatGPT and GPT-4 translate radiology reports into accurate and patient-friendly plain language?
- RQ2What is the quality of ChatGPT and GPT-4 translated reports as rated by radiologists, in terms of missing or misinterpreted information?
- RQ3Do prompts and prompt optimization meaningfully improve translation quality and usefulness of generated suggestions?
- RQ4How do different prompting strategies (including ensemble methods) compare in translation performance?
- RQ5What are the limitations and potential safety considerations for clinical deployment?
Key findings
- ChatGPT translations achieved an average radiologist-rated score of 4.268 on a 5-point scale for reported reports.
- Average information missing was 0.080 points per chest CT and 0.066 per brain MRI; average incorrect information was 0.065 per translation.
- Overall, 76% of chest CT translations scored a 5 and 32% of brain MRI translations achieved a 5 (within reported ranges).
- GPT-4 translations significantly outperformed ChatGPT with the original prompt and with the optimized prompt, approaching near-perfect results in some conditions (e.g., 96.8% good with optimized prompt).
- Optimized prompts substantially improved completeness and reduced omissions and misinterpretations compared with vague prompts (e.g., good translation rose from 55.2% to 77.2%).
- About 37% of cases yielded specific, report-based suggestions for patients or providers; most suggestions were general and relevant (e.g., follow-up with doctors, communicate findings).
- Prompt engineering and ensemble methods provided limited, non-significant gains over optimized prompts in many scenarios; ensembles sometimes introduced oversimplification or minor omissions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.