Skip to main content
QUICK REVIEW

[Paper Review] Comprehensive Evaluation of Multimodal AI Models in Medical Imaging Diagnosis: From Data Augmentation to Preference-Based Comparison

Cailian Ruan, Chengyue Huang|arXiv (Cornell University)|Dec 7, 2024
Radiomics and Machine Learning in Medical Imaging8 citations
TL;DR

The paper presents an evaluation framework for multimodal medical imaging diagnostics, expanding a CT-case dataset and comparing general-purpose multimodal models against vision-focused models and physicians using a preference-based assessment.

ABSTRACT

This study introduces an evaluation framework for multimodal models in medical imaging diagnostics. We developed a pipeline incorporating data preprocessing, model inference, and preference-based evaluation, expanding an initial set of 500 clinical cases to 3,000 through controlled augmentation. Our method combined medical images with clinical observations to generate assessments, using Claude 3.5 Sonnet for independent evaluation against physician-authored diagnoses. The results indicated varying performance across models, with Llama 3.2-90B outperforming human diagnoses in 85.27% of cases. In contrast, specialized vision models like BLIP2 and Llava showed preferences in 41.36% and 46.77% of cases, respectively. This framework highlights the potential of large multimodal models to outperform human diagnostics in certain tasks.

Motivation & Objective

  • Develop a standardized evaluation pipeline for multimodal models in abdominal CT diagnosis.
  • Augment a clinical dataset to enable robust model comparison.
  • Assess AI models against physician diagnoses using preference-based evaluation.
  • Contrast general-purpose vs. specialized vision models in complex diagnostic scenarios.

Proposed method

  • Preprocess data with de-identification, artifact handling, and synchronized image-text augmentation.
  • Encode 4-image CT sequences with paired diagnostic reports as standardized inputs for each model.
  • Use six multimodal models (four general-purpose, two specialized) to generate diagnostic reports.
  • Employ Claude 3.5 Sonnet as an independent assessor for three-way AI superior / physician superior / equivalent preferences.
  • Apply chi-square tests with Bonferroni correction to compare preference rates across models.
Figure 1: Comparative Evaluation Framework for Multimodal Medical Diagnosis
Figure 1: Comparative Evaluation Framework for Multimodal Medical Diagnosis

Experimental results

Research questions

  • RQ1Can general-purpose multimodal models outperform physicians in complex abdominal CT diagnoses?
  • RQ2How do specialized vision models compare to general-purpose models in multi-structure diagnostic tasks?
  • RQ3Does the proposed preference-based evaluation reliably distinguish AI and human diagnostic capabilities?

Key findings

  • General-purpose models outperform physician diagnoses in most cases, with Llama 3.2-90B achieving 85.27% AI Superior.
  • GPT-4, GPT-4o, and Gemini-1.5 also show high AI Superior rates (83.08%, 81.72%, 79.35%).
  • Specialized vision models BLIP2 and Llava achieve lower AI Superior rates (41.36% and 46.77%).
  • Equivalence rates are low (around 1.39%) across most models, indicating clear performance differences.
  • Statistical tests show p-values < 0.001 for general-purpose models; BLIP2 and Llava show p = 0.047 and p = 0.052 respectively.
Figure 2: Comparative Analysis Between Human and AI Diagnostic Assessments
Figure 2: Comparative Analysis Between Human and AI Diagnostic Assessments

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.