[Paper Review] Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks
Youku-mPLUG introduces the largest publicly available Chinese video-language pre-training dataset with 10 million high-quality video-text pairs collected from Youku, along with a comprehensive human-annotated benchmark for video-text retrieval, captioning, and classification. Pre-training on this dataset enables state-of-the-art performance, including a 23.1% accuracy gain in video category classification and 80.5% top-1 accuracy with the mPLUG-video model.
To promote the development of Vision-Language Pre-training (VLP) and multimodal Large Language Model (LLM) in the Chinese community, we firstly release the largest public Chinese high-quality video-language dataset named Youku-mPLUG, which is collected from Youku, a well-known Chinese video-sharing website, with strict criteria of safety, diversity, and quality. Youku-mPLUG contains 10 million Chinese video-text pairs filtered from 400 million raw videos across a wide range of 45 diverse categories for large-scale pre-training. In addition, to facilitate a comprehensive evaluation of video-language models, we carefully build the largest human-annotated Chinese benchmarks covering three popular video-language tasks of cross-modal retrieval, video captioning, and video category classification. Youku-mPLUG can enable researchers to conduct more in-depth multimodal research and develop better applications in the future. Furthermore, we release popular video-language pre-training models, ALPRO and mPLUG-2, and our proposed modularized decoder-only model mPLUG-video pre-trained on Youku-mPLUG. Experiments show that models pre-trained on Youku-mPLUG gain up to 23.1% improvement in video category classification. Besides, mPLUG-video achieves a new state-of-the-art result on these benchmarks with 80.5% top-1 accuracy in video category classification and 68.9 CIDEr score in video captioning, respectively. Finally, we scale up mPLUG-video based on the frozen Bloomz with only 1.7% trainable parameters as Chinese multimodal LLM, and demonstrate impressive instruction and video understanding ability. The zero-shot instruction understanding experiment indicates that pretraining with Youku-mPLUG can enhance the ability to comprehend overall and detailed visual semantics, recognize scene text, and leverage open-domain knowledge.
Motivation & Objective
- To address the lack of large-scale, high-quality public Chinese video-language datasets that hinder progress in vision-language pre-training (VLP) and multimodal LLM development.
- To overcome the absence of standardized, publicly available benchmarks for evaluating video-language models in Chinese, which currently limits fair comparison and reproducibility.
- To enable more effective pre-training and downstream application development in Chinese multimodal AI by providing a diverse, safe, and high-quality dataset with strict filtering and data cleaning.
- To establish a foundation for training advanced multimodal models, including a novel modularized decoder-only architecture (mPLUG-video), pre-trained on Youku-mPLUG.
- To demonstrate the effectiveness of pre-training on Chinese-specific data through significant performance gains in zero-shot instruction understanding and downstream tasks.
Proposed method
- The dataset is collected from 400 million raw videos on Youku, a major Chinese video-sharing platform, with strict filtering for safety, diversity, and quality using multi-level risk detection and manual curation.
- 10 million video-text pairs are selected from 45 diverse categories, ensuring balanced distribution and high data quality through text and video-level cleaning and a Chinese image-text pre-trained model.
- A human-annotated benchmark of 365,000 videos is constructed for three tasks: cross-modal retrieval, video captioning, and video category classification, enabling comprehensive evaluation.
- A new modularized decoder-only model, mPLUG-video, is proposed and pre-trained on Youku-mPLUG to enhance multimodal understanding and reasoning.
- The model is further scaled up using a frozen Bloomz backbone with only 1.7% trainable parameters to create a Chinese multimodal LLM, enabling strong zero-shot instruction following.
- Evaluation includes zero-shot instruction understanding with human-annotated cases, comparing mPLUG-video against VideoLLaMA and mPLUG-Video without pretraining.
Experimental results
Research questions
- RQ1Can a large-scale, high-quality Chinese video-language dataset significantly improve the performance of video-language models in downstream tasks?
- RQ2To what extent does pre-training on Youku-mPLUG enhance zero-shot video instruction understanding, particularly in recognizing visual semantics, scene text, and open-domain knowledge?
- RQ3How does modality complementarity between vision and language affect retrieval performance in video-language models?
- RQ4Can a decoder-only architecture pre-trained on Youku-mPLUG achieve state-of-the-art results in video category classification and captioning?
- RQ5Does pre-training on a culturally and linguistically specific dataset like Youku-mPLUG lead to better generalization and understanding of Chinese visual and linguistic concepts?
Key findings
- Models pre-trained on Youku-mPLUG achieve a 23.1% improvement in video category classification accuracy compared to baseline models.
- The mPLUG-video model achieves a new state-of-the-art result with 80.5% top-1 accuracy on video category classification and 68.9 CIDEr score on video captioning.
- Human evaluation shows that mPLUG-video outperforms VideoLLaMA and non-pretrained mPLUG-Video in zero-shot instruction understanding, with a higher proportion of correct and satisfying responses (A-level).
- The model demonstrates superior ability to understand fine-grained visual semantics, such as actions like 'jumping' and 'twisting', and accurately recognize scene text in videos.
- mPLUG-video shows enhanced capability in leveraging open-domain knowledge, correctly identifying key characters like 'Ultraman' in video contexts where other models fail.
- Pre-training with Youku-mPLUG significantly improves the model’s ability to comprehend overall video semantics, detailed visual cues, and linguistic nuances in Chinese.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.