Skip to main content
QUICK REVIEW

[논문 리뷰] Multimodal Foundation Models Exploit Text to Make Medical Image Predictions

Thomas Buckley, James A. Diao|arXiv (Cornell University)|2023. 11. 09.
Artificial Intelligence in Healthcare and Education인용 수 16
한 줄 요약

이 논문은 멀티모달 의료 AI 모델이 의료 이미지를 기반으로 예측할 때 주로 텍스트 정보에 의존하며, 텍스트가 이미지 기반 성능을 개선할 수도 있고 크게 악화시킬 수도 있음을 보여준다.

ABSTRACT

Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text, remain poorly understood. Here, using a diverse collection of 1014 multimodal medical cases, we evaluate the unimodal and multimodal image interpretation abilities of proprietary (GPT-4, Gemini Pro 1.0) and open-source (Llama-3.2-90B, LLaVA-Med-v1.5) multimodal foundational models with and without the use of text descriptions. Across all models, image predictions were largely driven by exploiting text, with accuracy increasing monotonically with the amount of informative text. By contrast, human performance on medical image interpretation did not improve with informative text. Exploitation of text is a double-edged sword; we show that even mild suggestions of an incorrect diagnosis in text diminishes image-based classification, reducing performance dramatically in cases the model could previously answer with images alone. Finally, we conducted a physician evaluation of model performance on long-form medical cases, finding that the provision of images either reduced or had no effect on model performance when text is already highly informative. Our results suggest that multimodal AI models may be useful in medical diagnostic reasoning but that their accuracy is largely driven, for better and worse, by their exploitation of text.

연구 동기 및 목표

  • 텍스트 설명이 있거나 없는 상태에서 멀티모달 기초 모델이 의료 이미지를 해석하는 방식을 평가한다.
  • 독점 및 오픈소스 모델 전반에서 이미지와 텍스트의 상대적 기여를 검토한다.
  • 정보성이 높은 텍스트가 모델 예측과 의학 영상 해석에 대한 인간 성능에 미치는 영향을 조사한다.
  • 경미한 부정확한 텍스트 프롬프트가 이미지 기반 분류 결과에 미치는 영향을 평가한다.
  • 장문 의료 사례에서 모델 성능에 대한 의사의 관점을 제공한다.

제안 방법

  • 다양한 1014개의 멀티모달 의료 사례를 모은다.
  • 모델들(GPT-4, Gemini Pro, Llama-3.2-90B, LLaVA-Med-v1.5) 간 단일모달 및 멀티모달 영상 해석을 평가한다.
  • 텍스트 설명 여부에 따른 모델 성능을 비교한다.
  • 텍스트의 서술성이 예측 정확도와 어떻게 상관하는지 분석한다.
  • 장문 사례에서 텍스트/이미지 입력과 이미지 전용 입력을 비교하기 위한 의사 평가를 수행한다.

실험 결과

연구 질문

  • RQ1멀티모달 모델이 의료 예측에서 이미지 데이터보다 텍스트에 더 의존하는가?
  • RQ2텍스트의 양과 정보성이 다양한 모델 전반에서 정확도에 어떤 영향을 미치는가?
  • RQ3텍스트 정보를 제공하는 것이 의학 영상에 대한 인간과 동등한 해석을 개선하는가 아니면 악화시키는가?
  • RQ4부정확하거나 오해의 소지가 있는 텍스트 프롬프트의 도입이 이미지 기반 예측에 미치는 효과는 무엇인가?
  • RQ5의사 평가에 따라 장문 임상 상황에서 텍스트와 이미지는 어떻게 상호작용하는가?

주요 결과

  • 모델 전반에서의 이미지 예측은 주로 텍스트를 악용하는 데서 비롯됐다.
  • 더 정보성이 높은 텍스트로 정확도가 모델 전반에서 증가했다.
  • 의료 영상 해석에 대한 인간의 성능은 정보성이 높은 텍스트로 향상되지 않았다.
  • 경미한 부정확한 텍스트 프롬프트는 이미지 기반 분류 성능을 상당히 저하시킬 수 있다.
  • 장문 사례에서 매우 정보성이 높은 텍스트를 포함한 이미지를 제시하면 모델 성능이 감소하거나 향상되지 않았다.
  • 결과는 멀티모달 모델이 진단 추론을 돕는 데 가능성을 시사하지만 텍스트 활용에 크게 좌우됨을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.