[论文解读] Comprehensive Evaluation of Multimodal AI Models in Medical Imaging Diagnosis: From Data Augmentation to Preference-Based Comparison
本文提出了一个多模态医学影像诊断的评估框架,扩展了 CT病例数据集,并在偏好基准评估下将通用多模态模型与以视觉为重点的模型以及医生进行对比。
This study introduces an evaluation framework for multimodal models in medical imaging diagnostics. We developed a pipeline incorporating data preprocessing, model inference, and preference-based evaluation, expanding an initial set of 500 clinical cases to 3,000 through controlled augmentation. Our method combined medical images with clinical observations to generate assessments, using Claude 3.5 Sonnet for independent evaluation against physician-authored diagnoses. The results indicated varying performance across models, with Llama 3.2-90B outperforming human diagnoses in 85.27% of cases. In contrast, specialized vision models like BLIP2 and Llava showed preferences in 41.36% and 46.77% of cases, respectively. This framework highlights the potential of large multimodal models to outperform human diagnostics in certain tasks.
研究动机与目标
- 为腹部 CT 诊断中的多模态模型开发标准化评估流程。
- 扩充临床数据集以实现鲁棒的模型比较。
- 使用基于偏好的评估方法,将 AI 模型与医生诊断进行对比。
- 在复杂诊断场景中对比通用型与专用视觉模型。
提出的方法
- 通过去识别、伪影处理和同步的图像-文本增强来预处理数据。
- 将含有诊断报告配对的 4 图像 CT 序列编码为每个模型的标准化输入。
- 使用六种多模态模型(四种通用、两种专用)生成诊断报告。
- 将 Claude 3.5 Sonnet 作为独立评估者,用于三方偏好:AI 优势/医生优势/等效。
- 对各模型偏好率进行卡方检验并进行 Bonferroni 纠正以比较。

实验结果
研究问题
- RQ1通用型多模态模型在复杂腹部 CT 诊断中能否优于医生?
- RQ2专用视觉模型在多结构诊断任务中与通用模型相比如何?
- RQ3所提出的基于偏好的评估能否可靠地区分 AI 与人类诊断能力?
主要发现
- 在大多数情况下,通用型模型的诊断优于医生,Llama 3.2-90B 达到 85.27% 的 AI 优势。
- GPT-4、GPT-4o 和 Gemini-1.5 也显示出较高的 AI 优势率(83.08%、81.72%、79.35%)。
- 专用视觉模型 BLIP2 和 Llava 的 AI 优势率较低(41.36% 和 46.77%)。
- 等效率在大多数模型中较低(约 1.39%),表明存在明显的性能差异。
- 统计检验证明通用模型的 p 值 < 0.001;BLIP2 和 Llava 的 p 值分别为 0.047 和 0.052。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。