[Paper Review] Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
The paper shows that multimodal medical AI models largely rely on textual information to predict from medical images, and that text can both improve and dramatically degrade image-based performance.
Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text, remain poorly understood. Here, using a diverse collection of 1014 multimodal medical cases, we evaluate the unimodal and multimodal image interpretation abilities of proprietary (GPT-4, Gemini Pro 1.0) and open-source (Llama-3.2-90B, LLaVA-Med-v1.5) multimodal foundational models with and without the use of text descriptions. Across all models, image predictions were largely driven by exploiting text, with accuracy increasing monotonically with the amount of informative text. By contrast, human performance on medical image interpretation did not improve with informative text. Exploitation of text is a double-edged sword; we show that even mild suggestions of an incorrect diagnosis in text diminishes image-based classification, reducing performance dramatically in cases the model could previously answer with images alone. Finally, we conducted a physician evaluation of model performance on long-form medical cases, finding that the provision of images either reduced or had no effect on model performance when text is already highly informative. Our results suggest that multimodal AI models may be useful in medical diagnostic reasoning but that their accuracy is largely driven, for better and worse, by their exploitation of text.
Motivation & Objective
- Assess how multimodal foundation models interpret medical images with and without text descriptions.
- Examine the relative contribution of images and text across proprietary and open-source models.
- Investigate how informative text affects model predictions and human performance on medical image interpretation.
- Evaluate the impact of mild incorrect textual prompts on image-based classification outcomes.
- Provide physician perspectives on model performance in long-form medical cases.
Proposed method
- Assemble a diverse set of 1014 multimodal medical cases.
- Evaluate unimodal and multimodal image interpretation across models (GPT-4, Gemini Pro, Llama-3.2-90B, LLaVA-Med-v1.5).
- Compare model performance with and without text descriptions.
- Analyze how text descriptiveness correlates with prediction accuracy.
- Conduct a physician evaluation on long-form cases to compare image-only vs text/ image inputs.
Experimental results
Research questions
- RQ1Do multimodal models rely more on text than on image data for medical predictions?
- RQ2How does the amount and informativeness of text impact model accuracy across different models?
- RQ3Does providing text information improve or degrade human-equivalent interpretation of medical images?
- RQ4What is the effect of introducing incorrect or misleading text prompts on image-based predictions?
- RQ5How do text and images interact in long-form clinical scenarios according to physician evaluation?
Key findings
- Image predictions across models were largely driven by exploiting text.
- Accuracy increased with more informative text across models.
- Human performance on medical image interpretation did not improve with informative text.
- Mild incorrect text prompts can substantially diminish image-based classification performance.
- In long-form cases, providing images with highly informative text either reduced or did not improve model performance.
- Results suggest multimodal models may aid diagnostic reasoning but are heavily influenced by text exploitation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.