Skip to main content
QUICK REVIEW

[논문 리뷰] GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

Pengcheng Chen, Ye Jin|arXiv (Cornell University)|2024. 08. 06.
Artificial Intelligence in Healthcare인용 수 8
한 줄 요약

GMAI-MMBench는 LVLM을 평가하기 위해 39개 모달리티에 걸친 285개 데이터셋을 포함하는 포괄적 다중모달 의료 AI 벤치마크로, GPT-4o와 같은 최상위 모델조차 대략 52%의 정확도를 달성한다는 점을 보고하고, 주요 미비점을 강조합니다.

ABSTRACT

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is crucial to develop benchmarks to evaluate LVLMs' effectiveness in various medical applications. Current benchmarks are often built upon specific academic literature, mainly focusing on a single domain, and lacking varying perceptual granularities. Thus, they face specific challenges, including limited clinical relevance, incomplete evaluations, and insufficient guidance for interactive LVLMs. To address these limitations, we developed the GMAI-MMBench, the most comprehensive general medical AI benchmark with well-categorized data structure and multi-perceptual granularity to date. It is constructed from 284 datasets across 38 medical image modalities, 18 clinical-related tasks, 18 departments, and 4 perceptual granularities in a Visual Question Answering (VQA) format. Additionally, we implemented a lexical tree structure that allows users to customize evaluation tasks, accommodating various assessment needs and substantially supporting medical AI research and applications. We evaluated 50 LVLMs, and the results show that even the advanced GPT-4o only achieves an accuracy of 53.96%, indicating significant room for improvement. Moreover, we identified five key insufficiencies in current cutting-edge LVLMs that need to be addressed to advance the development of better medical applications. We believe that GMAI-MMBench will stimulate the community to build the next generation of LVLMs toward GMAI.

연구 동기 및 목표

  • 다양한 모달리티, 작업 및 부서에 적용 가능한 일반 의료 AI(GMAI)에 대해 포괄적이며 임상적으로 관련된 다중모달 벤치마크를 구축한다.
  • 특정 임상 필요에 맞춘 맞춤형 평가를 가능하게 하는 잘 구성된 어휘 트리(렉시컬 트리)를 제공한다.
  • 의료 특화 LVLM, 오픈 소스 및 독점 모델을 포함한 다양한 LVLM을 평가하여 의료 AI의 강점, 약점 및 개선 영역을 규명한다.
  • 실제 임상 시나리오에서 대화형 LVLM의 지각적 세분화 요구사항(이미지, 영역 수준)에 대한 통찰을 제공한다.

제안 방법

  • 공개 소스와 병원에서 수집한 285개의 고품질 데이터셋을 모아 39개 모달리티와 18개 임상 VQA 작업을 18개 부서에 걸쳐 포괄한다.
  • 일관성을 보장하고 모호성을 줄이기 위해 SA-Med2D-20M 프로토콜과 MeSH 용어를 사용하여 이미지와 라벨을 표준화한다.
  • 맞춤형 평가를 가능하게 하려면 18개의 임상 VQA 작업, 18개의 부서, 4가지 지각적 세분성으로 구성된 렉시컬 트리를 구축한다.
  • 모달리티, 작업 큐, 및 세분성 주석을 포함한 26K개의 QA 쌍을 생성하고 품질과 균형을 위해 수동 검증과 선정을 수행한다.
  • VLMEvalKit 및 Multi-Modality-Arena 프레임워크를 사용하여 0샷 설정에서 44개의 LVLM(오픈 소스 및 의료 특화)과 6개의 독점 모델을 평가한다.

실험 결과

연구 질문

  • RQ1현행 LVLM이 광범위하고 임상적으로 현실적인 의료 모달리티와 과제에서 어떤 성능을 보이는가?
  • RQ2이미지, 영역(box), 마스크, 등 contour 등 다양한 지각적 세분성과 인터랙티브 큐로 평가했을 때 LVLM의 성능은 어떻게 달라지는가?
  • RQ3의료 진단과 추론에서 LVLM의 한계를 가장 많이 제한하는 요인은 무엇이며, 의료 도메인 모델은 일반 목적 모델과 어떻게 비교되는가?
  • RQ4잘 분류된 렉시컬 트리 벤치마크가 다양한 임상 부서와 요구에 대해 맞춤형 평가를 지원할 수 있는가?

주요 결과

  • GPT-4o는 GMAI-MMBench에서 52.24% 정확도를 달성하여 임상 과제에서 상당한 개선 여지가 있음을 시사한다.
  • MedDr 및 DeepSeek-VL-7B와 같은 오픈 소스 LVLM은 약 41%의 정확도에 도달하여 일부 독점 모델과 비교해도 경쟁력 있는 성능을 보인다.
  • 대부분의 의료 특화 LVLM은 중간 수준의 성능(~30%)에 도달하기 어렵고, 일부 설정에서 MedDr가 의료 특화 모델 중 최고를 달성한다.
  • 박스 수준의 지각은 이미지 수준이나 다른 세분성에 비해 일관되게 가장 낮은 정확도를 보이며, 영역 기반 추론의 문제를 강조한다.
  • 주요 병목은 지각 오류, 제한된 의료 도메인 지식, 무관한 응답, 그리고 답변의 안전성/거부 등이다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.