[Paper Review] Video Quality Assessment: A Comprehensive Survey
This paper provides a comprehensive survey of video quality assessment (VQA) methods, covering traditional, no-reference, and learning-based approaches. It systematically reviews metrics, models, datasets, and evaluation frameworks, offering a unified overview of VQA research and identifying key challenges and future directions in the field.
Video quality assessment (VQA) is an important processing task, aiming at predicting the quality of videos in a manner highly consistent with human judgments of perceived quality. Traditional VQA models based on natural image and/or video statistics, which are inspired both by models of projected images of the real world and by dual models of the human visual system, deliver only limited prediction performances on real-world user-generated content (UGC), as exemplified in recent large-scale VQA databases containing large numbers of diverse video contents crawled from the web. Fortunately, recent advances in deep neural networks and Large Multimodality Models (LMMs) have enabled significant progress in solving this problem, yielding better results than prior handcrafted models. Numerous deep learning-based VQA models have been developed, with progress in this direction driven by the creation of content-diverse, large-scale human-labeled databases that supply ground truth psychometric video quality data. Here, we present a comprehensive survey of recent progress in the development of VQA algorithms and the benchmarking studies and databases that make them possible. We also analyze open research directions on study design and VQA algorithm architectures. Github link: https://github.com/taco-group/Video-Quality-Assessment-A-Comprehensive-Survey.
Motivation & Objective
- To provide a systematic and up-to-date review of video quality assessment (VQA) techniques.
- To analyze the evolution and current state of no-reference, reduced-reference, and full-reference VQA models.
- To evaluate the performance and limitations of existing VQA metrics using standardized datasets.
- To identify open challenges and future research directions in VQA, including deep learning integration and perceptual modeling.
- To serve as a reference for researchers and practitioners in video quality research and applications.
Proposed method
- Categorize VQA methods into full-reference, reduced-reference, and no-reference frameworks based on available reference information.
- Review classical metrics such as PSNR, SSIM, and VMAF, and analyze their perceptual relevance and limitations.
- Examine learning-based VQA models that leverage convolutional neural networks (CNNs) and metric learning for quality prediction.
- Analyze the role of perceptual quality models, including structural and statistical features, in improving prediction accuracy.
- Survey widely used benchmark datasets such as TID2013, LIVE, and KoNViD-1k, and evaluate their diversity and limitations.
- Compare evaluation protocols and metrics used in VQA research, including correlation coefficients (e.g., SROCC, LCC) and statistical significance tests.
Experimental results
Research questions
- RQ1How do different VQA models (full-reference, reduced-reference, no-reference) compare in terms of accuracy and robustness?
- RQ2What are the key limitations of traditional VQA metrics like PSNR and SSIM in capturing human visual perception?
- RQ3To what extent do deep learning-based VQA models outperform classical approaches in predicting subjective video quality?
- RQ4How do benchmark datasets influence the generalization and evaluation of VQA models?
- RQ5What are the open challenges in developing perceptually accurate, scalable, and generalizable VQA solutions?
Key findings
- Traditional metrics such as PSNR and SSIM show limited correlation with human opinion due to their reliance on pixel-level differences.
- No-reference VQA models based on deep learning achieve higher correlation with subjective mean opinion scores (MOS) than classical methods.
- The VMAF metric demonstrates strong performance on streaming and compressed video, especially in adaptive bitrate scenarios.
- Datasets like KoNViD-1k and TID2013 are widely used but exhibit biases in content type, distortion types, and quality range.
- The integration of perceptual features and attention mechanisms in deep learning models improves robustness across diverse video content.
- Despite progress, generalization across different video codecs, resolutions, and content types remains a significant challenge in VQA.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.