[论文解读] The YLI-MED Corpus: Characteristics, Procedures, and Plans
YLI-MED语料库是一个公开可用的、人工标注的数据集,包含来自YFCC100M的50,700段用户生成视频,其中2,000段视频被标注为10种目标多媒体事件(例如,生日派对、婚礼),其余48,700段为非事件视频。该语料库提供详细的标注信息,包括事件类型、标注者一致性、置信度分数,以及语言和音乐等非事件属性,支持多媒体事件检测系统的训练与评估,并提供标准化的训练/测试划分及用于音频、视觉和运动分析的特征包。
The YLI Multimedia Event Detection corpus is a public-domain index of videos with annotations and computed features, specialized for research in multimedia event detection (MED), i.e., automatically identifying what's happening in a video by analyzing the audio and visual content. The videos indexed in the YLI-MED corpus are a subset of the larger YLI feature corpus, which is being developed by the International Computer Science Institute and Lawrence Livermore National Laboratory based on the Yahoo Flickr Creative Commons 100 Million (YFCC100M) dataset. The videos in YLI-MED are categorized as depicting one of ten target events, or no target event, and are annotated for additional attributes like language spoken and whether the video has a musical score. The annotations also include degree of annotator agreement and average annotator confidence scores for the event categorization of each video. Version 1.0 of YLI-MED includes 1823 "positive" videos that depict the target events and 48,138 "negative" videos, as well as 177 supplementary videos that are similar to event videos but are not positive examples. Our goal in producing YLI-MED is to be as open about our data and procedures as possible. This report describes the procedures used to collect the corpus; gives detailed descriptive statistics about the corpus makeup (and how video attributes affected annotators' judgments); discusses possible biases in the corpus introduced by our procedural choices and compares it with the most similar existing dataset, TRECVID MED's HAVIC corpus; and gives an overview of our future plans for expanding the annotation effort.
研究动机与目标
- 创建一个大规模、公开可访问的用户生成视频语料库,标注多媒体事件,以支持自动化事件检测的研究。
- 通过多阶段验证、共识评分和多位标注者的置信度估计,确保高质量的标注。
- 通过提供标准化的训练集和测试集,实现系统间的公平比较,确保事件分布均衡并包含负样本。
- 在YFCC100M数据集基础上,增加事件及非事件特征(如语言和音乐)的丰富结构化标注。
- 为未来的多媒体基因组计划(MMGP)奠定基础,拓展事件覆盖范围,并对视频内容进行更广泛的标注。
提出的方法
- 通过迭代研究和共识制定事件定义,确保标注者能够清晰理解并保持一致性。
- 使用关键词和元数据过滤从YFCC100M中收集视频,随后进行人工审查和事件及非事件特征的标注。
- 利用标注者间一致性度量和主观置信度评分,计算标注一致性与置信度分数,以评估标签的可靠性。
- 在移除低一致性视频并基于共识进行过滤以减少用户偏差后,将语料库划分为训练集和测试集。
- 通过自动化元数据过滤选择负样本视频,并经人工验证,确保其不包含任何目标事件。
- 预先计算并发布音频嵌入(YFCC-CC)、视觉特征(AlexNet)和图像特征(LIRE)等特征包,与标注信息一同发布。
实验结果
研究问题
- RQ1如何系统性地构建一个大规模、公开可用的视频语料库,用于多媒体事件检测,并确保高标注质量?
- RQ2非事件特征(如语言和音乐)在多大程度上影响事件检测性能?
- RQ3事件类型分布和标注者一致性如何影响事件检测模型的可靠性与泛化能力?
- RQ4数据收集过程引入了哪些主要偏差,以及在语料库设计中如何加以缓解?
- RQ5能否以该语料库为基础,有效启动一个可扩展的协作式大规模视频标注框架(如多媒体基因组计划)?
主要发现
- YLI-MED v.1.0语料库包含10种目标事件的2,000段正样本视频和48,700段非事件视频,训练集与测试集划分均衡。
- 事件分类的标注者一致性处于中等到较高水平,平均置信度分数反映出标注者具有较高的可靠性。
- 后期制作效果(如音乐和剪辑)显著影响事件检测决策,其中音乐是显著的混淆因素。
- 语言特征(包括非英文语音)在相当比例的视频中存在,并与事件分类结果相关。
- 该语料库设计上可与TRECVID MED和HAVIC相比较,但因流程差异,可能影响不同数据集间的直接可比性。
- 音频、视觉和运动特征的特征包已预先计算并发布,可直接用于下游事件检测研究。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。