Skip to main content
QUICK REVIEW

[Paper Review] Holistic Evaluation of GPT-4V for Biomedical Imaging

Zhengliang Liu, Hanqi Jiang|arXiv (Cornell University)|Nov 10, 2023
Artificial Intelligence in Healthcare and Education12 citations
TL;DR

This paper conducts a large-scale, multi-domain assessment of GPT-4V’s abilities on biomedical imaging tasks, including modality recognition, localization, diagnosis, and report generation across 16 imaging domains. It identifies strengths in modality/anatomy recognition and image captioning, and limitations in disease diagnosis and precise localization.

ABSTRACT

In this paper, we present a large-scale evaluation probing GPT-4V's capabilities and limitations for biomedical image analysis. GPT-4V represents a breakthrough in artificial general intelligence (AGI) for computer vision, with applications in the biomedical domain. We assess GPT-4V's performance across 16 medical imaging categories, including radiology, oncology, ophthalmology, pathology, and more. Tasks include modality recognition, anatomy localization, disease diagnosis, report generation, and lesion detection. The extensive experiments provide insights into GPT-4V's strengths and weaknesses. Results show GPT-4V's proficiency in modality and anatomy recognition but difficulty with disease diagnosis and localization. GPT-4V excels at diagnostic report generation, indicating strong image captioning skills. While promising for biomedical imaging AI, GPT-4V requires further enhancement and validation before clinical deployment. We emphasize responsible development and testing for trustworthy integration of biomedical AGI. This rigorous evaluation of GPT-4V on diverse medical images advances understanding of multimodal large language models (LLMs) and guides future work toward impactful healthcare applications.

Motivation & Objective

  • Evaluate GPT-4V across diverse biomedical imaging modalities (e.g., X-ray, MRI, CT, microscopy) to assess modality recognition capabilities.
  • Assess GPT-4V’s ability to localize anatomical structures within biomedical images.
  • Benchmark GPT-4V’s image classification/diagnosis performance on biomedical tasks.
  • Test GPT-4V’s capacity to generate diagnostic-style reports from medical images.
  • Provide insights into strengths, limitations, and implications for responsible, clinical deployment of biomedical AGI.

Proposed method

  • Utilize GPT-4V to perform zero-shot and direct classification on multiple public chest radiography datasets (e.g., MIMIC-CXR, CheXpert, ChestXray2017, COVID-Qu-Ex, OpenI, SIIM-ACR, NIH Chest X-rays).
  • Evaluate GPT-4V’s ability to output multi-class labels and to generate structured diagnostic reports with findings and impressions.
  • Extend evaluation to neuroimaging, oncological imaging for radiotherapy planning, cytopathology, ophthalmology, medical robotics, neurological disease imaging, biological imaging, cardiac imaging, ultrasound, nuclear medicine, endoscopy, dermatology, genetics, orthopedic/pediatric, and dental imaging.
  • In neuroimaging, test on high-resolution datasets (e.g., R1741 mouse brain, Allen Brain Institute atlas) to assess knowledge transfer and reconstruction quality evaluation.
  • In oncological imaging, use TCIA datasets (Burdenko-GBM-Progression, GLIS-RT, Lung-PET-CT-Dx) and breast X-ray (DDSM) for multimodal assessment.
  • For cytopathology, employ LC25000 and ALL datasets to evaluate cellular-level diagnostic capabilities, including limitations in staging.

Experimental results

Research questions

  • RQ1Can GPT-4V reliably recognize imaging modalities and anatomical regions across diverse biomedical datasets?
  • RQ2How well does GPT-4V perform disease diagnosis and lesion localization in biomedical images?
  • RQ3What is GPT-4V’s ability to generate coherent and clinically useful radiology or pathology reports from images?
  • RQ4What are the strengths and limitations of GPT-4V when applied to non-radiology biomedical imaging tasks (e.g., ophthalmology, pathology, neuroscience, biology, robotics)?
  • RQ5What considerations are necessary for responsible validation and deployment of biomedical multimodal LLMs like GPT-4V in clinical or research settings?

Key findings

  • GPT-4V shows proficiency in modality recognition and anatomy localization across multiple imaging domains.
  • GPT-4V excels at diagnostic report generation, indicating strong image captioning capabilities.
  • GPT-4V experiences difficulty with disease diagnosis and precise localization of lesions in some tasks.
  • Performance improvements are observed with additional contextual prompts and domain information, but gaps remain for high-stakes clinical decisions.
  • The study emphasizes the need for rigorous validation, bias assessment, and human-in-the-loop oversight for biomedical AGI deployment.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.