Skip to main content
QUICK REVIEW

[Paper Review] Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

Gunnar A. Sigurdsson, Abhinav Gupta|arXiv (Cornell University)|Apr 25, 2018
Video Surveillance and Tracking Methods8 references97 citations
TL;DR

Introduces Charades-Ego, a large-scale dataset with paired third- and first-person videos for joint egocentric and third-person action understanding, including 68,536 activity instances over 68.8 hours and accompanying annotations.

ABSTRACT

In Actor and Observer we introduced a dataset linking the first and third-person video understanding domains, the Charades-Ego Dataset. In this paper we describe the egocentric aspect of the dataset and present annotations for Charades-Ego with 68,536 activity instances in 68.8 hours of first and third-person video, making it one of the largest and most diverse egocentric datasets available. Charades-Ego furthermore shares activity classes, scripts, and methodology with the Charades dataset, that consist of additional 82.3 hours of third-person video with 66,500 activity instances. Charades-Ego has temporal annotations and textual descriptions, making it suitable for egocentric video classification, localization, captioning, and new tasks utilizing the cross-modal nature of the data.

Motivation & Objective

  • Link third- and first-person video understanding to leverage abundant third-person data for egocentric understanding.
  • Provide a large, diverse egocentric dataset with paired views and temporal/text annotations.
  • Enable tasks in egocentric video classification, localization, and captioning using cross-modal data.

Proposed method

  • Collect paired first- and third-person videos by having workers record scripts in both views using forehead-mounted and standard third-person setups.
  • Annotate videos temporally with cross-view annotations by presenting paired videos to workers and using the third-person view to inform first-person annotations.
  • Share activity classes, scripts, and methodology with the Charades dataset to ensure consistency across domains.
  • Split data into training and test sets with no subject overlap (80/20 split).
  • Provide baselines trained on both first- and third-person data and evaluate cross-view transfer and zero-shot egocentric recognition.

Experimental results

Research questions

  • RQ1Can first- and third-person videos be jointly learned to improve egocentric understanding?
  • RQ2How well do models trained on third-person data transfer to egocentric video tasks?
  • RQ3What is the benefit of incorporating first-person annotations for egocentric video classification and localization?

Key findings

  • Charades-Ego contains 68,536 egocentric activity instances in 68.8 hours of paired video (first and third person).
  • An additional 82.3 hours of third-person video with 66,500 activity instances from Charades are shared with Charades-Ego in methodology and classes.
  • First-person labelled data improves performance on egocentric test sets when training on Charades-Ego data (e.g., First/Third training achieves 28.2 in Table 1 for First-Person Test).
  • Zero-shot egocentric baselines show competitive performance (e.g., ResNet-152 Charades-Ego achieves 28.2 on egocentric test).
  • Using third-person trained models alone does not improve third-person accuracy when a strong third-person model exists; incorporating first-person data yields gains in egocentric tasks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.