[论文解读] Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
该论文提出 Hollywood in Homes 方法,通过众包实现端到端视频制作与注释,用于日常活动,形成 Charades 数据集,包含 9,848 个来源于家庭活动的视频及丰富注释,并提供基线评估用于行为识别与描述生成。
Computer vision has a great potential to help our daily lives by searching for lost keys, watering flowers or reminding us to take a pill. To succeed with such tasks, computer vision methods need to be trained from real and diverse examples of our daily dynamic scenes. While most of such scenes are not particularly exciting, they typically do not appear on YouTube, in movies or TV broadcasts. So how do we collect sufficiently many diverse but boring samples representing our lives? We propose a novel Hollywood in Homes approach to collect such data. Instead of shooting videos in the lab, we ensure diversity by distributing and crowdsourcing the whole process of video creation from script writing to video recording and annotation. Following this procedure we collect a new dataset, Charades, with hundreds of people recording videos in their own homes, acting out casual everyday activities. The dataset is composed of 9,848 annotated videos with an average length of 30 seconds, showing activities of 267 people from three continents. Each video is annotated by multiple free-text descriptions, action labels, action intervals and classes of interacted objects. In total, Charades provides 27,847 video descriptions, 66,500 temporally localized intervals for 157 action classes and 41,104 labels for 46 object classes. Using this rich data, we evaluate and provide baseline results for several tasks including action recognition and automatic description generation. We believe that the realism, diversity, and casual nature of this dataset will present unique challenges and new opportunities for computer vision community.
研究动机与目标
- 激发对真实、多样、日常生活数据的需求,超越 YouTube/电影和实验室记录。
- 提出一个覆盖脚本编写、拍摄和注释的众包数据收集流程,以捕捉乏味的日常活动。
- 创建一个具有丰富时间性动作和对象交互注释的大规模、多样化数据集(Charades)。
- 为 Charades 提供基线评估,用于动作识别和自动描述生成。
提出的方法
- 使用包含 40 个对象和 30 个动作的词汇表来引导基于场景的提示进行众包脚本生成。
- 众包视频录制,工人在家中按剧本情句表演约 30 秒。
- 众包验证与注释,包括 157 个动作类别的时间定位和对象交互,以及自由文本描述。
- 三阶段 AMT 工作流:脚本生成、视频拍摄和注释/验证。
- 训练与评估划分的构造,防止训练与测试之间的工人重叠,并平衡类别分布。
实验结果
研究问题
- RQ1众包、有剧本的家庭内视频是否能够提供现实且多样的日常活动数据,超越娱乐视频?
- RQ2使用标准方法和最先进方法,在 Charades 上进行动作识别和描述生成的基线性能水平是多少?
- RQ3与不受控的在线视频相比,对象-动作交互和场景上下文在一个受控词汇表的众包数据集中如何表现?
主要发现
- Charades 包含 9,848 个视频(平均 30.1s),在 157 个动作类别中共有 66,500 个时间定位的动作区间。
- 数据集包含 46 个对象类别和 30 个动词词汇,促成出现的动作-对象交互。
- 基线动作识别使用改进轨迹、CNN 基于方法和双流方法 yields 相对温和的 mAP,其中 IDT 在 17.2% mAP 处表现最好,Combined 达到 18.6%。
- 句子预测显示 S2VT 作为生成描述的最强基线,CIDEr 分数指示相比人类描述仍有提升空间。
- 数据揭示动作共现和情境丰富的交互,反映真实世界的日常活动,突出细粒度动作识别和视频字幕的挑战。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。