[Paper Review] Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
This paper introduces MATH-Vision (MATH-V), a new benchmark of 3,040 high-quality mathematical reasoning problems with visual contexts drawn from real math competitions. Curated across 16 subjects and 5 difficulty levels, it reveals a significant performance gap—LMMs achieve 22.76% vs. human 75.66%—highlighting limitations in current models' multimodal mathematical reasoning and enabling detailed error analysis for future improvement.
Recent advancements in Large Multimodal Models (LMMs) have shown promising results in mathematical reasoning within visual contexts, with models approaching human-level performance on existing benchmarks such as MathVista. However, we observe significant limitations in the diversity of questions and breadth of subjects covered by these benchmarks. To address this issue, we present the MATH-Vision (MATH-V) dataset, a meticulously curated collection of 3,040 high-quality mathematical problems with visual contexts sourced from real math competitions. Spanning 16 distinct mathematical disciplines and graded across 5 levels of difficulty, our dataset provides a comprehensive and diverse set of challenges for evaluating the mathematical reasoning abilities of LMMs. Through extensive experimentation, we unveil a notable performance gap between current LMMs and human performance on MATH-V, underscoring the imperative for further advancements in LMMs. Moreover, our detailed categorization allows for a thorough error analysis of LMMs, offering valuable insights to guide future research and development. The project is available at https://mathvision-cuhk.github.io
Motivation & Objective
- To address the lack of diversity and breadth in existing benchmarks for multimodal mathematical reasoning, such as MathVista, which suffer from limited question types and narrow subject coverage.
- To develop a comprehensive, high-quality dataset of math problems with visual contexts derived from real math competitions, ensuring accuracy and consistency through expert validation.
- To enable fine-grained evaluation of Large Multimodal Models (LMMs) across diverse mathematical disciplines and difficulty levels, facilitating detailed error analysis.
- To establish a more rigorous and representative testbed to assess the true mathematical reasoning capabilities of LMMs beyond current benchmarks.
Proposed method
- The MATH-Vision dataset is constructed from 19 real math competitions, ensuring problems are authentic, challenging, and representative of advanced mathematical reasoning.
- All 3,040 problems are manually curated and cross-validated by expert annotators to guarantee unique, correct answers and high data quality.
- Problems are categorized into 16 distinct mathematical disciplines—such as logic, arithmetic, graph theory, analytic geometry, and combinatorial geometry—using a human-verified classification system.
- The dataset includes 1,532 open-ended and 1,508 multiple-choice questions, ensuring balanced evaluation formats.
- Each problem is assigned a difficulty level from 1 (easiest) to 5 (most challenging), enabling granular performance analysis across difficulty tiers.
- Extensive experiments are conducted on four leading LMMs (GPT-4V, Gemini, SPHINX, InternLM-XComposer2-VL) to evaluate performance across subjects and difficulty levels.
Experimental results
Research questions
- RQ1To what extent do current Large Multimodal Models (LMMs) generalize across diverse mathematical domains when presented with visual math problems from real competitions?
- RQ2How does model performance vary across different difficulty levels and mathematical subjects in visual reasoning tasks?
- RQ3What are the specific failure modes of LMMs in multimodal mathematical reasoning, and how can they be diagnosed through structured error analysis?
- RQ4How does the performance of LMMs trained on internal data (e.g., GPT-4V, Gemini) compare to those trained on public data (e.g., SPHINX) on a high-quality, competition-based benchmark?
- RQ5Can a curated, expert-validated benchmark like MATH-Vision reveal meaningful performance gaps between LMMs and humans that are obscured by existing, less diverse benchmarks?
Key findings
- Current LMMs achieve only 22.76% accuracy on the MATH-Vision benchmark, while human performance reaches 75.66%, revealing a substantial 53% performance gap.
- The performance gap is most pronounced in advanced mathematical domains such as topology, metric geometry, and combinatorial geometry, where models struggle with complex spatial and abstract reasoning.
- GPT-4V and Gemini, models trained on internal data, outperform public-data-trained models like SPHINX, indicating the impact of data quality and scale on reasoning capability.
- Error analysis shows that models frequently fail to derive correct geometric relationships or misinterpret visual configurations, especially in problems requiring multi-step reasoning or non-trivial symmetry insights.
- The dataset’s 16-subject and 5-level difficulty categorization enables precise diagnosis of model weaknesses, such as poor performance in logic or solid geometry despite strong performance in basic arithmetic.
- The benchmark exposes limitations in current LMMs’ ability to generalize beyond pattern-matching on simple visual cues, underscoring the need for deeper reasoning mechanisms.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.