Skip to main content
QUICK REVIEW

[논문 리뷰] ChatRadio-Valuer: A Chat Large Language Model for Generalizable Radiology Report Generation Based on Multi-institution and Multi-system Data

Tianyang Zhong, Wei Zhao|arXiv (Cornell University)|2023. 10. 08.
Artificial Intelligence in Healthcare and Education인용 수 12
한 줄 요약

ChatRadio-Valuer는 332,673개의 영상의학 보고서를 다기관 및 다계통 데이터에서 Llama2를 미세조정하여 영상의학 보고서 생성의 일반화를 도모하고, 영상의학 보고서로부터 질병 진단에서 ChatGPT 및 GPT-4를 능가합니다.

ABSTRACT

Radiology report generation, as a key step in medical image analysis, is critical to the quantitative analysis of clinically informed decision-making levels. However, complex and diverse radiology reports with cross-source heterogeneity pose a huge generalizability challenge to the current methods under massive data volume, mainly because the style and normativity of radiology reports are obviously distinctive among institutions, body regions inspected and radiologists. Recently, the advent of large language models (LLM) offers great potential for recognizing signs of health conditions. To resolve the above problem, we collaborate with the Second Xiangya Hospital in China and propose ChatRadio-Valuer based on the LLM, a tailored model for automatic radiology report generation that learns generalizable representations and provides a basis pattern for model adaptation in sophisticated analysts' cases. Specifically, ChatRadio-Valuer is trained based on the radiology reports from a single institution by means of supervised fine-tuning, and then adapted to disease diagnosis tasks for human multi-system evaluation (i.e., chest, abdomen, muscle-skeleton, head, and maxillofacial $\&$ neck) from six different institutions in clinical-level events. The clinical dataset utilized in this study encompasses a remarkable total of extbf{332,673} observations. From the comprehensive results on engineering indicators, clinical efficacy and deployment cost metrics, it can be shown that ChatRadio-Valuer consistently outperforms state-of-the-art models, especially ChatGPT (GPT-3.5-Turbo) and GPT-4 et al., in terms of the diseases diagnosis from radiology reports. ChatRadio-Valuer provides an effective avenue to boost model generalization performance and alleviate the annotation workload of experts to enable the promotion of clinical AI applications in radiology reports.

연구 동기 및 목표

  • 다수의 기관과 신체 시스템에 걸쳐 일반화 가능한 완전하고 임상적으로 실행 가능한 영상의학 보고서 생성 솔루션을 개발한다.
  • 단일 기관 미세조정을 사용하여 기관 간 적응 영상의학 보고서 생성을 가능하게 한다.
  • 여섯 기관과 다섯 신체 시스템 전반의 일반화 능력을 평가한다.
  • 실제 임상 AI 도입을 촉진하기 위한 임상 유용성 및 배포 비용을 평가한다.

제안 방법

  • 일반화 가능한 영상의학 지식을 학습하기 위해 대규모 영상의학 보고서 코퍼스에서 Llama2를 미세조정한다.
  • 전문가 주도 청소, 프롬프트 합성, 다계통/다기관 통합을 통해 고품질 프롬프트를 생성하도록 데이터 전처리한다.
  • 학습/평가 80/20 분할을Construct하며 Institution 1의 데이터를 미세조정에 사용하고 others를 테스트에 사용한다.
  • 발견 사항을 입력으로 LLM의 impression을 추출하여 영상의학 보고서 소견을 생성한다.
  • 엔지니어링 지표 및 전문가 주도 임상 유용성 평가를 사용하여 최첨단 모델과 비교한다.
Figure 3 : The architecture diagram of Llama 2. The model structure of Llama 2 is basically consistent with the standard Transformer Decoder structure, mainly composed of 32 Transformer Blocks
Figure 3 : The architecture diagram of Llama 2. The model structure of Llama 2 is basically consistent with the standard Transformer Decoder structure, mainly composed of 32 Transformer Blocks

실험 결과

연구 질문

  • RQ1ChatRadio-Valuer가 여섯 기관과 다섯 영상의학 시스템 전반에 걸쳐 기관 간 일반화를 달성할 수 있는가?
  • RQ2,
  • RQ3ChatRadio-Valuer는 영상의학 보고서 생성 및 보고서에서의 질병 진단에서 최첨단 모델(ChatGPT, GPT-4 등)과 어떻게 비교되는가?
  • RQ4이 접근 방식이 주석 작업 부하 및 실용적 임상 유용성에 어떤 영향을 미치는가?
  • RQ5이상 이질적인 영상의학 데이터에서 robust 일반화를 가능하게 하는 데이터 전처리 및 프롬프트 전략은 무엇인가?

주요 결과

  • ChatRadio-Valuer는 영상의학 보고서에서 질병 진단에 있어 최첨단 모델을 지속적으로 능가한다.
  • 프레임워크는 여섯 기관 및 다섯 시스템에서 기관 간 및 다계통 일반화를 입증한다.
  • 데이터 전처리 및 전문가가 선별한 프롬프트는 노이즈를 줄이고 프롬프트 품질을 향상시켜 견고한 미세조정을 가능하게 한다.
  • 이 접근 방식은 임상 효능 평가 및 배포 비용 고려를 지원하여 실세계 영상의학 AI 도입에 기여한다.
  • 모델은 컨텍스트 길이 4096, FFN의 SwiGLU, RMSNorm, RoPE, 그리고 그룹드-쿼리 어텐션을 활용하여 이질적인 영상의학 데이터를 처리한다.
Figure 4 : Prompt generation overview. The overall framework contains three parts, system description, instruction, and input, which collaboratively constitute a prompt. Within a prompt example (purple), expert instruction and input data on its right are inserted to the { Expert Instruction } and {
Figure 4 : Prompt generation overview. The overall framework contains three parts, system description, instruction, and input, which collaboratively constitute a prompt. Within a prompt example (purple), expert instruction and input data on its right are inserted to the { Expert Instruction } and {

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.