Skip to main content
QUICK REVIEW

[Paper Review] A Survey on Benchmarks of Multimodal Large Language Models

Li Jian, Weisheng Lu|arXiv (Cornell University)|Aug 16, 2024
Topic Modeling4 citations
TL;DR

This survey presents a comprehensive analysis of 180 benchmarks for Multimodal Large Language Models (MLLMs), categorizing them into perception and understanding, cognition and reasoning, specific domains, key capabilities, and other modalities. It identifies GPT-4 and Gemini as top-performing models across 83 benchmarks since 2024 and argues that evaluation must be treated as a core discipline to advance MLLM development and AGI.

ABSTRACT

Multimodal Large Language Models (MLLMs) are gaining increasing popularity in both academia and industry due to their remarkable performance in various applications such as visual question answering, visual perception, understanding, and reasoning. Over the past few years, significant efforts have been made to examine MLLMs from multiple perspectives. This paper presents a comprehensive review of 200 benchmarks and evaluations for MLLMs, focusing on (1)perception and understanding, (2)cognition and reasoning, (3)specific domains, (4)key capabilities, and (5)other modalities. Finally, we discuss the limitations of the current evaluation methods for MLLMs and explore promising future directions. Our key argument is that evaluation should be regarded as a crucial discipline to support the development of MLLMs better. For more details, please visit our GitHub repository: https://github.com/swordlidev/Evaluation-Multimodal-LLMs-Survey.

Motivation & Objective

  • To provide a comprehensive, systematic review of 180 benchmarks evaluating Multimodal Large Language Models (MLLMs).
  • To identify and categorize evaluation benchmarks along five key dimensions: perception and understanding, cognition and reasoning, specific domains, key capabilities, and other modalities.
  • To analyze the performance of top MLLMs (e.g., GPT-4, Gemini) across 83 benchmarks since 2024 to reveal trends and model strengths.
  • To highlight limitations in current evaluation practices and advocate for evaluation as a central discipline in MLLM development.
  • To guide future research by identifying under-explored areas and promoting robust, reliable, and generalizable evaluation frameworks.

Proposed method

  • Systematic collection and classification of 180 MLLM evaluation benchmarks into five primary categories based on evaluation focus.
  • Curation of benchmark details including task types, data size, modality support (image, video, audio, 3D, etc.), annotation format, and evaluation protocols.
  • Performance analysis of the top three MLLMs (GPT-4, Gemini, GPT-4V) on 83 benchmarks from 2024, including metrics like accuracy and task coverage.
  • Taxonomy development to organize benchmarks by functional scope (e.g., spatial reasoning, hallucination detection, medical vision), enabling cross-benchmark comparison.
  • Inclusion of emerging benchmarks such as M3DBench (3D reasoning), ScanReason (3D grounding), and MMT-Bench (omnimodal tasks) to reflect evolving evaluation needs.
  • Use of GPT-based evaluation for open-ended benchmarks (e.g., MQA, open-ended questions) to standardize comparison across models.

Experimental results

Research questions

  • RQ1What are the dominant evaluation categories and sub-types in MLLM benchmarking, and how are they distributed across perception, reasoning, and domain-specific tasks?
  • RQ2How do top-performing MLLMs (e.g., GPT-4, Gemini) compare in performance across 83 benchmarks published since 2024?
  • RQ3What are the key limitations in current MLLM evaluation practices, particularly regarding robustness, fairness, and real-world generalization?
  • RQ4How do benchmarks targeting emerging modalities (e.g., 3D, audio, point clouds) contribute to evaluating advanced MLLM capabilities?
  • RQ5To what extent do existing benchmarks assess critical user-facing capabilities such as instruction following, hallucination avoidance, and multi-turn reasoning?

Key findings

  • GPT-4 and Gemini consistently outperformed other models across 83 benchmarks since 2024, indicating their dominance in multimodal reasoning and understanding.
  • Perception and understanding benchmarks (e.g., VQA, image captioning) remain the most prevalent, while cognition and reasoning tasks (e.g., spatial reasoning, logical inference) are growing in complexity.
  • Benchmarks targeting 3D environments (e.g., M3DBench, ScanReason) and spatial reasoning (e.g., SpatialRGPT) are emerging to evaluate fine-grained spatial understanding, a current weakness in MLLMs.
  • Only 15% of benchmarks evaluated hallucination or trustworthiness, highlighting a critical gap in assessing model reliability and safety.
  • The MMT-Bench and MCUB benchmarks represent a shift toward evaluating multimodal commonality and joint reasoning across four modalities (image, audio, video, point cloud), signaling a move toward holistic multimodal intelligence evaluation.
  • Despite the proliferation of benchmarks, no unified evaluation framework exists, and many benchmarks lack standardization in annotation, evaluation metrics, or open access.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.