Skip to main content
QUICK REVIEW

[论文解读] A Survey on Machine Learning Techniques for Auto Labeling of Video, Audio, and Text Data

Shikun Zhang, Omid Jafari|arXiv (Cornell University)|Sep 8, 2021
Music and Audio Processing参考文献 89被引用 27
一句话总结

对跨视频、音频和文本数据的优化数据标注与标注方法的综述,涵盖无监督、半监督、监督、主动学习和迁移学习策略,以及标注工具。

ABSTRACT

Machine learning has been utilized to perform tasks in many different domains such as classification, object detection, image segmentation and natural language analysis. Data labeling has always been one of the most important tasks in machine learning. However, labeling large amounts of data increases the monetary cost in machine learning. As a result, researchers started to focus on reducing data annotation and labeling costs. Transfer learning was designed and widely used as an efficient approach that can reasonably reduce the negative impact of limited data, which in turn, reduces the data preparation cost. Even transferring previous knowledge from a source domain reduces the amount of data needed in a target domain. However, large amounts of annotated data are still demanded to build robust models and improve the prediction accuracy of the model. Therefore, researchers started to pay more attention on auto annotation and labeling. In this survey paper, we provide a review of previous techniques that focuses on optimized data annotation and labeling for video, audio, and text data.

研究动机与目标

  • 解释监督学习中数据标注的重要性与成本,以及自动标注的动机。
  • 对跨视频、音频和文本领域的优化标注方法进行调查和分类。
  • 总结现有支持自动或半自动标注的标注工具和框架。
  • 突出与以图像为中心的综述的差异,并指出未来研究的空白。

提出的方法

  • 回顾关于视频数据的自动和半自动标注技术的文献,包括无监督、半监督、监督、主动学习、迁移学习和多标签方法。
  • 回顾音频数据标注的文献,重点是无监督、半监督、监督、主动学习和多标签方法。
  • 回顾文本数据标注的文献,聚焦于命名实体识别、文本分类和词性标注及自动/半自动策略。
  • 总结可用于视频、音频和文本数据的标注工具并讨论它们的功能与部署环境。
  • 将发现组织成有结构的分类法并提出未来工作的方向。

实验结果

研究问题

  • RQ1有哪些用于优化视频、音频和文本数据标注的主要策略?
  • RQ2无监督、半监督、监督、主动学习和迁移学习方法在不同领域的比较如何?
  • RQ3存在哪些标注工具,它们如何支持自动或半自动标注?
  • RQ4提出的未来方向是什么,以进一步降低标注成本并提高鲁棒性?

主要发现

  • 视频数据优化按学习范式进行分类(无监督、半监督、监督、主动学习、迁移学习),并按多标签和图结构方法进行区分。
  • 音频数据标注利用无监督特征学习、半监督与监督标注、用于分段/说话人识别的主动学习,以及用于捕捉标签相关性的多标签方法。
  • 文本数据标注的进展包括NER的预标注、半自动标注,以及用于提升准确性的领域特定分类法和嵌入策略。
  • 存在多种用于视频、音频和文本的实际标注工具,支持自动标注、半自动标注或基于云的标注服务。
  • 本综述强调将主动学习与迁移学习结合的潜力,并指出在深度强化学习和异构多模态数据方面的未来研究机会。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。