[Paper Review] A Survey of Multimodal Large Language Model from A Data-centric Perspective
This survey presents a comprehensive data-centric analysis of multimodal large language models (MLLMs), focusing on data collection, curation, quality, and evaluation across modalities. It proposes that iterative data improvement—through enhanced data quantity, quality, and structure—significantly boosts MLLM performance, offering a systematic framework for dataset development and benchmarking beyond model architecture optimization.
Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal role in the development and refinement of these models. In this survey, we comprehensively review the literature on MLLMs from a data-centric perspective. Specifically, we explore methods for preparing multimodal data during the pretraining and adaptation phases of MLLMs. Additionally, we analyze the evaluation methods for the datasets and review the benchmarks for evaluating MLLMs. Our survey also outlines potential future research directions. This work aims to provide researchers with a detailed understanding of the data-driven aspects of MLLMs, fostering further exploration and innovation in this field.
Motivation & Objective
- To address the lack of comprehensive studies on data-centric approaches in MLLMs, especially beyond model architecture improvements.
- To investigate how data characteristics—quantity, quality, and modality heterogeneity—affect MLLM performance during pretraining and adaptation.
- To provide a systematic review of existing multimodal datasets and evaluation benchmarks across vision, audio, text, and video modalities.
- To identify open challenges and propose future research directions in data curation, data management, and robust evaluation for MLLMs.
- To establish a foundation for data-driven innovation in MLLMs by emphasizing iterative data refinement over model-centric tuning.
Proposed method
- Systematically categorizes and reviews multimodal datasets across five modalities: text, vision, audio, video, and 3D environments.
- Analyzes data preparation pipelines for MLLMs, including data collection, selection, annotation, and preprocessing strategies during pretraining and fine-tuning.
- Evaluates data quality through metrics such as annotation consistency, modality alignment, and linguistic diversity.
- Reviews benchmarking frameworks for MLLMs, including task-specific datasets like VQA, VCR, Charades-STA, QVHighlight, AudioCaps, Clotho, and ClothoAQA.
- Proposes a data-centric framework emphasizing iterative data improvement—expanding data volume, enhancing data quality, and structuring data for better modality alignment.
- Integrates insights from data-centric AI (DCAI) to guide dataset design and evaluation, focusing on data as the primary driver of model performance.
Experimental results
Research questions
- RQ1How can multimodal data be effectively collected, selected, and managed across diverse modalities to support MLLM training?
- RQ2What is the impact of data quality, quantity, and modality alignment on MLLM performance across different tasks and architectures?
- RQ3How can existing evaluation benchmarks be improved to better reflect real-world multimodal understanding and reasoning capabilities?
- RQ4What are the key challenges in curating and maintaining high-quality, diverse, and scalable multimodal datasets for MLLMs?
- RQ5What future research directions are needed to advance data-centric development of MLLMs beyond current model-centric paradigms?
Key findings
- Data quality is as critical as data quantity; curated datasets can enable smaller models to match the performance of larger models.
- The Charades-STA dataset supports complex, long-form language queries for temporal activity localization, with 13,898 training and 4,233 test clip-sentence pairs.
- The QVHighlight dataset provides 10,000+ YouTube videos with query-relevant moment annotations and saliency scores, enabling highlight detection and moment localization.
- AudioCaps contains 46,000 audio-caption pairs, while Clotho offers 4,981 audio samples with 24,905 captions, supporting general audio captioning and description tasks.
- ClothoAQA provides 1,991 audio files with 11,946 crowdsourced questions (including yes/no and single-word answers), enabling audio question answering evaluation.
- Audio-MusicAVQA includes over 45,000 QA pairs from 9,000+ musical performance videos, supporting spatio-temporal reasoning in audio-visual scenes across 9 question types and 33 templates.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.