Skip to main content
QUICK REVIEW

[Paper Review] A Survey on Machine Learning Techniques for Auto Labeling of Video, Audio, and Text Data

Shikun Zhang, Omid Jafari|arXiv (Cornell University)|Sep 8, 2021
Music and Audio ProcessingComputer Science89 references27 citations
TL;DR

A survey of optimized data annotation and labeling methods across video, audio, and text data, covering unsupervised, semi-supervised, supervised, active, and transfer learning strategies as well as annotation tools.

ABSTRACT

Machine learning has been utilized to perform tasks in many different domains such as classification, object detection, image segmentation and natural language analysis. Data labeling has always been one of the most important tasks in machine learning. However, labeling large amounts of data increases the monetary cost in machine learning. As a result, researchers started to focus on reducing data annotation and labeling costs. Transfer learning was designed and widely used as an efficient approach that can reasonably reduce the negative impact of limited data, which in turn, reduces the data preparation cost. Even transferring previous knowledge from a source domain reduces the amount of data needed in a target domain. However, large amounts of annotated data are still demanded to build robust models and improve the prediction accuracy of the model. Therefore, researchers started to pay more attention on auto annotation and labeling. In this survey paper, we provide a review of previous techniques that focuses on optimized data annotation and labeling for video, audio, and text data.

Motivation & Objective

  • Explain the importance and cost of data labeling in supervised learning and the motivation for automated annotation.
  • Survey and categorize optimized annotation approaches across video, audio, and text domains.
  • Summarize existing annotation tools and frameworks that support automated or semi-automatic labeling.
  • Highlight differences from image-centric surveys and identify gaps for future research.

Proposed method

  • Review literature on automated and semi-automatic annotation techniques for video data, including unsupervised, semi-supervised, supervised, active learning, transfer learning, and multi-label approaches.
  • Review literature on audio data annotation with emphasis on unsupervised, semi-supervised, supervised, active learning, and multi-label methods.
  • Review literature on text data annotation focusing on named entity recognition, text classification, and part-of-speech tagging with automatic/semi-automatic strategies.
  • Summarize available annotation tools for video, audio, and text data and discuss their capabilities and deployment contexts.
  • Organize findings into a structured taxonomy and suggest avenues for future work.

Experimental results

Research questions

  • RQ1What are the main strategies used to optimize data annotation for video, audio, and text data?
  • RQ2How do unsupervised, semi-supervised, supervised, active learning, and transfer learning approaches compare across domains?
  • RQ3What annotation tools exist and how do they support automated or semi-automatic labeling?
  • RQ4What are the proposed future directions to further reduce labeling costs and improve robustness?

Key findings

  • Video data optimization is categorized by learning paradigm (unsupervised, semi-supervised, supervised, active learning, transfer learning) and by multi-label and graph-based methods.
  • Audio data annotation leverages unsupervised feature learning, semi-supervised and supervised tagging, active learning for diarization, and multi-label approaches to capture tag correlations.
  • Text data annotation advances include pre-annotation for NER, semi-automatic labeling, and domain-specific taxonomy and embedding strategies for improved accuracy.
  • Multiple practical annotation tools exist for video, audio, and text, enabling auto-annotation, semi-automatic labeling, or cloud-based annotation services.
  • The survey highlights the potential of combining active learning with transfer learning and notes opportunities in deep reinforcement learning and heterogeneous multi-modal data for future work.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.