[Paper Review] Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos
Introduces Charades-Ego, a large-scale dataset with paired third- and first-person videos for joint egocentric and third-person action understanding, including 68,536 activity instances over 68.8 hours and accompanying annotations.
In Actor and Observer we introduced a dataset linking the first and third-person video understanding domains, the Charades-Ego Dataset. In this paper we describe the egocentric aspect of the dataset and present annotations for Charades-Ego with 68,536 activity instances in 68.8 hours of first and third-person video, making it one of the largest and most diverse egocentric datasets available. Charades-Ego furthermore shares activity classes, scripts, and methodology with the Charades dataset, that consist of additional 82.3 hours of third-person video with 66,500 activity instances. Charades-Ego has temporal annotations and textual descriptions, making it suitable for egocentric video classification, localization, captioning, and new tasks utilizing the cross-modal nature of the data.
Motivation & Objective
- Link third- and first-person video understanding to leverage abundant third-person data for egocentric understanding.
- Provide a large, diverse egocentric dataset with paired views and temporal/text annotations.
- Enable tasks in egocentric video classification, localization, and captioning using cross-modal data.
Proposed method
- Collect paired first- and third-person videos by having workers record scripts in both views using forehead-mounted and standard third-person setups.
- Annotate videos temporally with cross-view annotations by presenting paired videos to workers and using the third-person view to inform first-person annotations.
- Share activity classes, scripts, and methodology with the Charades dataset to ensure consistency across domains.
- Split data into training and test sets with no subject overlap (80/20 split).
- Provide baselines trained on both first- and third-person data and evaluate cross-view transfer and zero-shot egocentric recognition.
Experimental results
Research questions
- RQ1Can first- and third-person videos be jointly learned to improve egocentric understanding?
- RQ2How well do models trained on third-person data transfer to egocentric video tasks?
- RQ3What is the benefit of incorporating first-person annotations for egocentric video classification and localization?
Key findings
- Charades-Ego contains 68,536 egocentric activity instances in 68.8 hours of paired video (first and third person).
- An additional 82.3 hours of third-person video with 66,500 activity instances from Charades are shared with Charades-Ego in methodology and classes.
- First-person labelled data improves performance on egocentric test sets when training on Charades-Ego data (e.g., First/Third training achieves 28.2 in Table 1 for First-Person Test).
- Zero-shot egocentric baselines show competitive performance (e.g., ResNet-152 Charades-Ego achieves 28.2 on egocentric test).
- Using third-person trained models alone does not improve third-person accuracy when a strong third-person model exists; incorporating first-person data yields gains in egocentric tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.