Skip to main content
QUICK REVIEW

[论文解读] TRECVID 2020: A comprehensive campaign for evaluating video retrieval tasks across multiple application domains

George Awad, Asad A. Butt|arXiv (Cornell University)|Apr 27, 2021
Multimodal Machine Learning Applications参考文献 20被引用 17
一句话总结

TRECVID 2020 对跨六个任务的视频检索开展了全面的评估活动,其中包括两个新任务:视频摘要和灾难场景描述与索引。该活动使用了多样化的数据集,如 Vimeo V3C1、BBC EastEnders、VIRAT 和尼泊尔地震视频,结合人工与自动评估,采用 MAP、MT 和 DA 等指标,标志着首次设立视频摘要任务,共有 12 支团队参与,其中 2 支团队提交了详细的解决方案。

ABSTRACT

The TREC Video Retrieval Evaluation (TRECVID) is a TREC-style video analysis and retrieval evaluation with the goal of promoting progress in research and development of content-based exploitation and retrieval of information from digital video via open, metrics-based evaluation. Over the last twenty years this effort has yielded a better understanding of how systems can effectively accomplish such processing and how one can reliably benchmark their performance. TRECVID has been funded by NIST (National Institute of Standards and Technology) and other US government agencies. In addition, many organizations and individuals worldwide contribute significant time and effort. TRECVID 2020 represented a continuation of four tasks and the addition of two new tasks. In total, 29 teams from various research organizations worldwide completed one or more of the following six tasks: 1. Ad-hoc Video Search (AVS), 2. Instance Search (INS), 3. Disaster Scene Description and Indexing (DSDI), 4. Video to Text Description (VTT), 5. Activities in Extended Video (ActEV), 6. Video Summarization (VSUM). This paper is an introduction to the evaluation framework, tasks, data, and measures used in the evaluation campaign.

研究动机与目标

  • 推动跨多个现实世界应用场景的内容驱动视频检索与分析。
  • 评估系统在即席搜索、实例搜索、视频摘要和长视频中的活动识别方面的表现。
  • 引入并评估新任务,如灾难场景描述与索引和视频摘要。
  • 提供标准化数据集和评估协议,以支持视频检索系统的可靠基准测试。
  • 通过开放、基于指标的评估(结合人工与自动评分)推动研究进展。

提出的方法

  • 在即席视频搜索和实例搜索任务中使用了 Vimeo Creative Commons V3C1 数据集(100 万段视频片段,约 1000 小时)。
  • 在实例搜索和视频摘要任务中使用了 BBC EastEnders 视频(464 小时),并采用 MTCNN 和 FaceNet 进行片段级分割和人脸检测。
  • 在长视频中的活动(ActEV)任务中使用了 VIRAT 数据集(10 小时),由 Kitware, Inc. 基于参考标注进行注释。
  • 在视频到文本描述任务中使用了 V3C2 的 1700 个视频子集,结合人工标注和通过 MT 指标与直接评估(DA)进行自动评分。
  • 引入了基于 2015 年尼泊尔地震公共灾难视频的 5 小时灾难场景描述与索引任务。
  • 实施混合评估:由人工评估员对 AVS、INS、VSUM 和 DSDI 任务进行评分;通过 MAP、MT 和 DA 自动评分 VTT 和 ActEV 任务。

实验结果

研究问题

  • RQ1系统如何在多样化的现实世界视频应用中有效检索相关视频内容?
  • RQ2在长篇叙事视频中识别重大生活事件以用于摘要时,面临哪些关键挑战?
  • RQ3系统在自动与人工评估下,能多大程度上从短片段生成准确且全面的视频描述?
  • RQ4视频分析系统在非结构化、现实世界的灾难视频中检测和索引事件的能力如何?
  • RQ5在视频检索与描述任务的评分中,不同评估方法(人工 vs. 自动)的比较结果如何?

主要发现

  • 视频摘要任务首次设立,共有 12 支团队参与,仅 2 支团队提交最终运行结果。
  • 两名获奖团队结合人脸检测得分与自注意力机制(如 VASNet)来计算视频片段的重要性,用于摘要生成。
  • 人工评估员对即席视频搜索、实例搜索和视频摘要任务进行评分;而自动指标(MAP、MT、DA)用于视频到文本任务和灾难场景任务。
  • ActEV 任务使用了 Kitware, Inc. 提供的参考标注,并采用标准评估协议进行评分。
  • DSDI 任务使用了 2015 年尼泊尔地震的 5 小时数据集,结果通过平均精度均值(MAP)进行评分。
  • 数据访问延迟可能对参与度造成了负面影响,尤其在新设立的视频摘要任务中。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。