Skip to main content
QUICK REVIEW

[논문 리뷰] A Transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics

Hong-Yu Zhou, Yizhou Yu|arXiv (Cornell University)|2023. 06. 01.
COVID-19 diagnosis using AI인용 수 11
한 줄 요약

트랜스포머 기반 모델이 멀티모달 임상 데이터를 이미지, 텍스트, 검사실 데이터, 인구통계까지 하나의 표현으로 통합하여 진단을 보조하고 COVID-19 결과를 예측하며, 이미지 전용 및 비통합 베이스라인보다 성능이 우수합니다.

ABSTRACT

During the diagnostic process, clinicians leverage multimodal information, such as chief complaints, medical images, and laboratory-test results. Deep-learning models for aiding diagnosis have yet to meet this requirement. Here we report a Transformer-based representation-learning model as a clinical diagnostic aid that processes multimodal input in a unified manner. Rather than learning modality-specific features, the model uses embedding layers to convert images and unstructured and structured text into visual tokens and text tokens, and bidirectional blocks with intramodal and intermodal attention to learn a holistic representation of radiographs, the unstructured chief complaint and clinical history, structured clinical information such as laboratory-test results and patient demographic information. The unified model outperformed an image-only model and non-unified multimodal diagnosis models in the identification of pulmonary diseases (by 12% and 9%, respectively) and in the prediction of adverse clinical outcomes in patients with COVID-19 (by 29% and 7%, respectively). Leveraging unified multimodal Transformer-based models may help streamline triage of patients and facilitate the clinical decision process.

연구 동기 및 목표

  • 진단 과정에서 멀티모달 임상 정보를 통합할 필요성을 제시한다.
  • 이미지, 비구조화 텍스트, 구조화된 데이터를 하나의 방식으로 처리하는 트랜스포머 기반 표현 학습 모델을 개발한다.
  • 이미지 전용 및 비통합 멀티모달 접근법보다 진단 성능의 개선을 보여준다.
  • triage와 임상 의사결정에 대한 잠재적 이점을 제시한다.
  • 방사선 영상, 주요 증상, 임상 과거력, 검사실 및 인구통계 정보에의 적용 가능성을 강조한다.

제안 방법

  • 이미지와 텍스트(비구조화 및 구조화된 텍스트) 를 시각적 토큰과 텍스트 토큰으로 변환하기 위해 임베딩 계층을 사용한다.
  • 내재 모달 및 교차 모달 주의력을 가진 양방향 트랜스포머 블록으로 전체적인 표현을 학습한다.
  • 방사선 영상, 주요 증상, 임상 과거력, 검사실 결과 및 인구통계 정보를 통합 아키텍처에서 처리한다.
  • 통합 멀티모달 모델과 이미지 전용 및 비통합 멀티모달 베이스라인을 비교한다.
  • 폐질환 식별 및 COVID-19 악성 결과 예측에 대해 평가한다.

실험 결과

연구 질문

  • RQ1통합 멀티모달 트랜스포머 모델이 이미지 전용 모델보다 폐질환 식별에서 더 우수할 수 있는가?
  • RQ2비통합 멀티모달 접근법과 비교하여 COVID-19의 악성 임상 결과 예측에서 통합 모델이 개선될 수 있는가?
  • RQ3여러 데이터 모달리티(이미지, 텍스트, 구조화된 데이터)를 통합하는 것이 진단 의사결정을 향상시키는가?
  • RQ4내재 모달 및 교차 모달 주의가 포괄적 임상 표현 학습에 어떤 영향을 미치는가?

주요 결과

  • 통합 멀티모달 모델은 폐질환 식별에서 이미지 전용 모델보다 12% 앞섰다.
  • 통합 모델은 비통합 멀티모달 모델보다 폐질환 식별에서 9% 앞섰다.
  • COVID-19 악성 결과에 대해, 통합 모델은 이미지 전용 베이스라인 대비 29%의 개선과 비통합 멀티모달 모델 대비 7%의 개선을 달성했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.