[论文解读] Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
本文表明多模态医学人工智能模型在很大程度上依赖文本信息来从医学影像中进行预测,文本既可以提升也能显著降低基于影像的性能。
Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text, remain poorly understood. Here, using a diverse collection of 1014 multimodal medical cases, we evaluate the unimodal and multimodal image interpretation abilities of proprietary (GPT-4, Gemini Pro 1.0) and open-source (Llama-3.2-90B, LLaVA-Med-v1.5) multimodal foundational models with and without the use of text descriptions. Across all models, image predictions were largely driven by exploiting text, with accuracy increasing monotonically with the amount of informative text. By contrast, human performance on medical image interpretation did not improve with informative text. Exploitation of text is a double-edged sword; we show that even mild suggestions of an incorrect diagnosis in text diminishes image-based classification, reducing performance dramatically in cases the model could previously answer with images alone. Finally, we conducted a physician evaluation of model performance on long-form medical cases, finding that the provision of images either reduced or had no effect on model performance when text is already highly informative. Our results suggest that multimodal AI models may be useful in medical diagnostic reasoning but that their accuracy is largely driven, for better and worse, by their exploitation of text.
研究动机与目标
- 评估多模态基础模型在有文本描述与无文本描述的情况下对医学影像的解读。
- 检查专有模型与开源模型中影像和文本的相对贡献。
- 研究有信息量的文本如何影响模型预测和人类在医学影像解读中的表现。
- 评估轻微错误文本提示对基于影像的分类结果的影响。
- 提供医生对在长篇医学病例中模型性能的看法。
提出的方法
- 汇集一组多元模态医学案例,总数为1014。
- 在不同模型(GPT-4、Gemini Pro、Llama-3.2-90B、LLaVA-Med-v1.5)中评估单模态和多模态影像解读。
- 比较有文本描述与无文本描述时的模型性能。
- 分析文本描述性与预测准确性之间的相关性。
- 在长篇病例中进行医生评估,以比较仅图像与文本/图像输入的差异。
实验结果
研究问题
- RQ1多模态模型在医疗预测中是否比对影像数据更依赖文本?
- RQ2文本的数量和信息量如何影响不同模型的准确性?
- RQ3提供文本信息是提升还是降低与人类相当的医学影像解读?
- RQ4引入不正确或具有误导性的文本提示对基于影像的预测有何影响?
- RQ5根据医生评估,文本与影像在长篇临床情境中的相互作用是怎样的?
主要发现
- 各模型的影像预测在很大程度上是通过利用文本来实现的。
- 在各模型中,随着文本信息量的增加,准确性提高。
- 人类在医学影像解读方面的表现并未因为信息丰富的文本而提升。
- 轻微错误的文本提示会显著降低基于影像的分类性能。
- 在长篇病例中,提供带有高度信息性文本的图像要么降低要么未能提升模型性能。
- 结果表明多模态模型可能有助于诊断推理,但高度依赖文本利用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。