Skip to main content
QUICK REVIEW

[论文解读] Space-Time Video Regularity and Visual Fidelity: Compression, Resolution and Frame Rate Adaptation

Dae Yeol Lee, Hyunsuk Ko|arXiv (Cornell University)|Mar 31, 2021
Image and Video Quality Assessment参考文献 52被引用 4
一句话总结

该论文提出了一种新颖的视频质量预测模型,利用时空位移帧差的自然视频统计(NVS)来预测在压缩、分辨率和帧率自适应组合下的感知质量。通过沿运动轨迹建模除法归一化的帧差,并使用学习回归器,该方法在ETRI-LIVE STSVQ数据库上实现了最先进性能,展现出对复杂视频失真的强鲁棒性。

ABSTRACT

In order to be able to deliver today's voluminous amount of video contents through limited bandwidth channels in a perceptually optimal way, it is important to consider perceptual trade-offs of compression and space-time downsampling protocols. In this direction, we have studied and developed new models of natural video statistics (NVS), which are useful because high-quality videos contain statistical regularities that are disturbed by distortions. Specifically, we model the statistics of divisively normalized difference between neighboring frames that are relatively displaced. In an extensive empirical study, we found that those paths of space-time displaced frame differences that provide maximal regularity against our NVS model generally align best with motion trajectories. Motivated by this, we build a new video quality prediction engine that extracts NVS features from displaced frame differences, and combines them in a learned regressor that can accurately predict perceptual quality. As a stringent test of the new model, we apply it to the difficult problem of predicting the quality of videos subjected not only to compression, but also to downsampling in space and/or time. We show that the new quality model achieves state-of-the-art (SOTA) prediction performance compared on the new ETRI-LIVE Space-Time Subsampled Video Quality (STSVQ) database, which is dedicated to this problem. Downsampling protocols are of high interest to the streaming video industry, given rapid increases in frame resolutions and frame rates.

研究动机与目标

  • 解决在压缩、空间下采样和时间下采样组合条件下预测感知视频质量的挑战。
  • 建模自然视频序列中因失真而被破坏的统计规律性。
  • 开发一种利用时空位移帧差的质量预测引擎,以提升保真度估计的准确性。
  • 在专为时空下采样设计的新颖且具有挑战性的基准数据集上评估该模型。
  • 为带宽受限的流媒体环境提供一种感知准确且可学习的视频质量评估框架。

提出的方法

  • 该方法建模时间与空间位移视频帧之间的除法归一化差异,以捕捉时空规律性。
  • 识别在学习到的自然视频统计(NVS)模型下使规律性最大化的位移帧差路径。
  • 从这些位移帧差中提取基于NVS的特征,作为学习回归器的输入。
  • 使用包含多样化失真的大规模数据集,训练回归器以预测主观视频质量得分。
  • 该方法在ETRI-LIVE空间-时间下采样视频质量(STSVQ)数据库上进行了验证,该数据库包含压缩、分辨率和帧率变化的组合。
  • 使用运动轨迹对齐作为感知质量的代理,增强了对真实失真的敏感性。

实验结果

研究问题

  • RQ1如何利用时空位移帧差来建模自然视频统计以实现质量预测?
  • RQ2在NVS模型下使规律性最大化的路径在多大程度上与实际运动轨迹对齐?
  • RQ3基于NVS特征的学习回归器能否在预测组合压缩、分辨率和帧率自适应下的质量方面超越现有模型?
  • RQ4该模型在专为时空下采样设计的基准数据集上的表现如何?
  • RQ5空间下采样与时间下采样对感知视频质量退化的影响各占多大比重?

主要发现

  • 所提出的模型在ETRI-LIVE STSVQ数据库上实现了最先进预测性能,在预测复杂现实失真下的质量方面优于现有方法。
  • 在NVS模型下使规律性最大化的位移帧差路径与实际运动轨迹具有高度对齐,验证了模型的感知相关性。
  • 使用除法归一化的帧差显著提高了模型对感知失真的敏感性,相较于标准差分特征表现更优。
  • 学习到的回归器能有效泛化于压缩、空间下采样和时间下采样各种组合。
  • 该模型对高失真水平表现出强鲁棒性,在极端带宽限制下仍保持高预测准确性。
  • 结果证实,时空规律性是视觉保真度的强预测因子,尤其在结合运动感知特征提取时更为显著。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。