[Paper Review] ChatRadio-Valuer: A Chat Large Language Model for Generalizable Radiology Report Generation Based on Multi-institution and Multi-system Data
ChatRadio-Valuer fine-tunes Llama2 on 332,673 radiology reports from multi-institution and multi-system data to generalize radiology report generation and outperform ChatGPT and GPT-4 in disease diagnosis from radiology reports.
Radiology report generation, as a key step in medical image analysis, is critical to the quantitative analysis of clinically informed decision-making levels. However, complex and diverse radiology reports with cross-source heterogeneity pose a huge generalizability challenge to the current methods under massive data volume, mainly because the style and normativity of radiology reports are obviously distinctive among institutions, body regions inspected and radiologists. Recently, the advent of large language models (LLM) offers great potential for recognizing signs of health conditions. To resolve the above problem, we collaborate with the Second Xiangya Hospital in China and propose ChatRadio-Valuer based on the LLM, a tailored model for automatic radiology report generation that learns generalizable representations and provides a basis pattern for model adaptation in sophisticated analysts' cases. Specifically, ChatRadio-Valuer is trained based on the radiology reports from a single institution by means of supervised fine-tuning, and then adapted to disease diagnosis tasks for human multi-system evaluation (i.e., chest, abdomen, muscle-skeleton, head, and maxillofacial $\&$ neck) from six different institutions in clinical-level events. The clinical dataset utilized in this study encompasses a remarkable total of extbf{332,673} observations. From the comprehensive results on engineering indicators, clinical efficacy and deployment cost metrics, it can be shown that ChatRadio-Valuer consistently outperforms state-of-the-art models, especially ChatGPT (GPT-3.5-Turbo) and GPT-4 et al., in terms of the diseases diagnosis from radiology reports. ChatRadio-Valuer provides an effective avenue to boost model generalization performance and alleviate the annotation workload of experts to enable the promotion of clinical AI applications in radiology reports.
Motivation & Objective
- Develop a complete, clinically viable radiology report generation solution that generalizes across multiple institutions and body systems.
- Enable cross-institution adaptive radiology report generation using single-institution fine-tuning samples.
- Evaluate generalization capabilities across six institutions and five body systems.
- Assess clinical utility and deployment costs to facilitate real-world radiology AI adoption.
Proposed method
- Fine-tune Llama2 on a large radiology report corpus to learn generalizable radiology knowledge.
- Preprocess data via expert-driven cleaning, prompt synthesis, and multi-system/institution consolidation to produce high-quality prompts.
- Construct an 80/20 train/evaluation split with data from Institution 1 used for fine-tuning and others for testing.
- Generate radiology report impressions by feeding findings to the LLM and extracting the impression.
- Evaluate using engineering metrics and expert-driven clinical utility assessments to compare with state-of-the-art models.

Experimental results
Research questions
- RQ1Can ChatRadio-Valuer achieve cross-institution generalization across six institutions and five radiology systems?
- RQ2How does ChatRadio-Valuer compare to state-of-the-art models (e.g., ChatGPT, GPT-4) in radiology report generation and disease diagnosis from reports?
- RQ3What impact does the approach have on annotation workload and practical clinical utility?
- RQ4What data preprocessing and prompting strategies are essential for enabling robust generalization across heterogeneous radiology data?
Key findings
- ChatRadio-Valuer consistently outperforms state-of-the-art models in diseases diagnosis from radiology reports.
- The framework demonstrates cross-institution and multi-system generalization on six institutions and five systems.
- Data preprocessing and expert-curated prompts reduce noise and improve prompt quality for robust fine-tuning.
- The approach supports clinical-efficacy evaluation and deployment-cost considerations, aiding real-world radiology AI deployment.
- The model leverages Llama2 architecture with context length 4096, SwiGLU in FFN, RMSNorm, RoPE, and grouped-query attention to handle heterogeneous radiology data.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.