Skip to main content
QUICK REVIEW

[Paper Review] Recent Standard Development Activities on Video Coding for Machines

Wen Gao, Shan Liu|arXiv (Cornell University)|May 26, 2021
Visual Attention and Saliency DetectionComputer Science7 references31 citations
TL;DR

This paper surveys MPEG VCM's recent activities, outlining use cases, requirements, processing pipelines, evaluation frameworks, and proposed technical solutions for video coding optimized for machine Vision tasks.

ABSTRACT

In recent years, video data has dominated internet traffic and becomes one of the major data formats. With the emerging 5G and internet of things (IoT) technologies, more and more videos are generated by edge devices, sent across networks, and consumed by machines. The volume of video consumed by machine is exceeding the volume of video consumed by humans. Machine vision tasks include object detection, segmentation, tracking, and other machine-based applications, which are quite different from those for human consumption. On the other hand, due to large volumes of video data, it is essential to compress video before transmission. Thus, efficient video coding for machines (VCM) has become an important topic in academia and industry. In July 2019, the international standardization organization, i.e., MPEG, created an Ad-Hoc group named VCM to study the requirements for potential standardization work. In this paper, we will address the recent development activities in the MPEG VCM group. Specifically, we will first provide an overview of the MPEG VCM group including use cases, requirements, processing pipelines, plan for potential VCM standards, followed by the evaluation framework including machine-vision tasks, dataset, evaluation metrics, and anchor generation. We then introduce technology solutions proposed so far and discuss the recent responses to the Call for Evidence issued by MPEG VCM group.

Motivation & Objective

  • Summarize the MPEG VCM group's scope, use cases, and requirements for video coding for machines.
  • Present the processing pipelines and planned standards for VCM.
  • Describe the evaluation framework including machine-vision tasks, datasets, metrics, and anchor generation.
  • Summarize the technology solutions proposed to date and responses to the Call for Evidence.

Proposed method

  • Review of MPEG VCM documentation and Call for Evidence responses.
  • Description of the VCM use cases, requirements, and processing pipelines.
  • Outline of the evaluation framework with machine-vision tasks, datasets, and metrics.
  • Overview of proposed technical solutions and their alignment with potential standardization.
  • Discussion of plan for potential VCM standards and anchor generation.

Experimental results

Research questions

  • RQ1What are the use cases and requirements driving Video Coding for Machines in MPEG VCM?
  • RQ2What processing pipelines are envisioned for VCM standardization?
  • RQ3What evaluation framework elements (tasks, datasets, metrics, anchors) are proposed for VCM?
  • RQ4What technological solutions have been proposed for VCM and how have they responded to the Call for Evidence?

Key findings

  • VCM addresses machine-centric video compression needs differing from human-centric use cases.
  • A structured evaluation framework with machine vision tasks, datasets, metrics, and anchor generation is proposed.
  • Technology solutions proposed so far are discussed in relation to MPEG VCM's Call for Evidence responses.
  • There is a clear plan for potential VCM standards and an outlined processing pipeline.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.