[論文レビュー] GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
GMAI-MMBench は、LVLMs を評価するための、39 種類のモダリティにわたる285のデータセットを備えた総合的なマルチモーダル医療AIベンチマークであり、GPT-4o のような最先端モデルですら約52%の精度を達成することを報告するとともに、主要な不備を浮き彫りにしています。
Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is crucial to develop benchmarks to evaluate LVLMs' effectiveness in various medical applications. Current benchmarks are often built upon specific academic literature, mainly focusing on a single domain, and lacking varying perceptual granularities. Thus, they face specific challenges, including limited clinical relevance, incomplete evaluations, and insufficient guidance for interactive LVLMs. To address these limitations, we developed the GMAI-MMBench, the most comprehensive general medical AI benchmark with well-categorized data structure and multi-perceptual granularity to date. It is constructed from 284 datasets across 38 medical image modalities, 18 clinical-related tasks, 18 departments, and 4 perceptual granularities in a Visual Question Answering (VQA) format. Additionally, we implemented a lexical tree structure that allows users to customize evaluation tasks, accommodating various assessment needs and substantially supporting medical AI research and applications. We evaluated 50 LVLMs, and the results show that even the advanced GPT-4o only achieves an accuracy of 53.96%, indicating significant room for improvement. Moreover, we identified five key insufficiencies in current cutting-edge LVLMs that need to be addressed to advance the development of better medical applications. We believe that GMAI-MMBench will stimulate the community to build the next generation of LVLMs toward GMAI.
研究の動機と目的
- モダリティ、タスク、部門を横断して適用可能な、総合的で臨床的に関連性の高いマルチモーダルベンチマーク(GMAI)を確立する。
- 特定の臨床ニーズに合わせて高度にカスタマイズ可能な評価を可能にする、構造化された語彙ツリーを提供する。
- 医療特化モデル、オープンソース、ライセンスモデルを含む幅広いLVLMを評価し、医療AIの強み・弱み・改善分野を特定する。
- 実務的な臨床シナリオにおける対話型LVLMの知覚的粒度要件(画像、領域レベル)に関する洞察を提供する。
提案手法
- 公開ソースと病院から、39モダリティ、18の臨床VQAタスク、18部門にまたがる285の高品質データセットを組み立てる。
- 一貫性を確保し、曖昧性を減らすため、SA-Med2D-20MプロトコルとMeSH用語を用いて画像とラベルを標準化する。
- 18の臨床VQAタスク、18の部門、4つの知覚的粒度を備えた語彙ツリーを構築して、カスタマイズ評価を可能にする。
- モダリティ、タスクキュー、粒度の注釈を付けた26KのQAペアを生成し、品質とバランスのための手動検証と選択を行う。
- VLMEvalKitとMulti-Modality-Arenaフレームワークを用いて、ゼロショット設定で44のLVLM(オープンソースおよび医療特化)と6つの専有モデルを評価する。
実験結果
リサーチクエスチョン
- RQ1臨床的に現実的な幅広い医療モダリティとタスクのセットに対する、現在のLVLMの性能はどの程度か。
- RQ2知覚的粒度(画像、ボックス、マスク、輪郭)や対話型キューを変えて評価したとき、LVLMはどのように機能するか。
- RQ3医療診断と推論を最も制限する要因は何か、医療ドメインモデルと一般モデルの比較はどうか。
- RQ4よくカテゴライズされた語彙ツーベンチマークは、さまざまな臨床部門とニーズに合わせたカスタマイズ評価を支援できるか。
主な発見
- GPT-4o は GMAI-MMBench で52.24%の精度を達成しており、臨床タスクの改善余地が大きいことを示している。
- MedDr や DeepSeek-VL-7B などのオープンソースLVLMは約41%の精度に達し、特定の専有モデルと比較して競争力のある性能を示している。
- 医療特化LVLMの大半は中間レベルの性能(約30%)に到達するのに苦戦しており、いくつかの設定でMedDrが医療特化モデルの中で最も良い成績を収めている。
- ボックスレベルの知覚は、画像レベルや他の粒度と比較して常に最も低い精度を示し、領域ベースの推論の課題を浮き彫りにしている。
- 主要なボトルネックには、知覚エラー、医療ドメイン知識の制限、無関係な回答、回答の安全性・拒否が含まれる。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。