[论文解读] Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos
介绍 Charades-Ego,这是一个具有配对的三人称和第一人称视频的大规模数据集,用于联合的自我中心与三人称动作理解,包含 68,536 个活动实例,总时长 68.8 小时及相应的注释。
In Actor and Observer we introduced a dataset linking the first and third-person video understanding domains, the Charades-Ego Dataset. In this paper we describe the egocentric aspect of the dataset and present annotations for Charades-Ego with 68,536 activity instances in 68.8 hours of first and third-person video, making it one of the largest and most diverse egocentric datasets available. Charades-Ego furthermore shares activity classes, scripts, and methodology with the Charades dataset, that consist of additional 82.3 hours of third-person video with 66,500 activity instances. Charades-Ego has temporal annotations and textual descriptions, making it suitable for egocentric video classification, localization, captioning, and new tasks utilizing the cross-modal nature of the data.
研究动机与目标
- 将第三人称和第一人称视频理解联系起来,以利用大量三人称数据来促进自我中心理解。
- 提供一个大规模、多样化的自我中心数据集,具有配对视角以及时间/文本注释。
- 使用跨模态数据实现自我中心视频分类、定位和字幕生成等任务的能力。
提出的方法
- 通过让工作人员在前额佩戴设备的视角和标准三人称设置下记录脚本,收集配对的第一人称与第三人称视频。
- 通过向工作人员展示配对视频,并使用第三人称视角来 informing 第一人称注释,对视频进行时间注释和跨视角注释。
- 与 Charades 数据集共享活动类别、脚本和方法,以确保跨领域的一致性。
- 将数据分为训练集与测试集,且不与受试者重叠(80/20 划分)。
- 提供在第一人称和第三人称数据上训练的基线,并评估跨视图迁移和零样本自我中心识别。
实验结果
研究问题
- RQ1第一人称和第三人称视频能否联合学习以提升自我中心理解?
- RQ2在第三人称数据上训练的模型在自我中心视频任务中的迁移效果如何?
- RQ3将第一人称注释纳入自我中心视频分类和定位的收益是什么?
主要发现
- Charades-Ego 在 68.8 小时的配对视频(第一人称和第三人称)中包含 68,536 个自我中心活动实例。
- 额外的 82.3 小时第三人称视频,包含 66,500 个来自 Charades 的活动实例,与 Charades-Ego 在方法学和类别上共同使用。
- 在 Charades-Ego 数据上进行训练时,第一人称标注数据提升了在自我中心测试集上的表现(例如 First/Third 训练在第一人称测试中的表 1 得分为 28.2)。
- 零样本自我中心基线显示出具有竞争力的表现(例如 ResNet-152 Charades 在自我中心测试中的分数为 28.2)。
- 仅使用第三人称训练的模型在存在强大第三人称模型时并不能提高三人称的准确性;将第一人称数据引入会在自我中心任务中带来增益。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。