Skip to main content
QUICK REVIEW

[Paper Review] EPIC-KITCHENS-100

Dima Damen, Davide Moltisanti|Explore Bristol Research|Jan 1, 2020
Video Surveillance and Tracking MethodsComputer Science129 references102 citations
TL;DR

EPIC-KITCHENS-100 extends the largest egocentric vision dataset to 100 hours of head-mounted video with denser, multi-task annotations, enabling action recognition, detection, anticipation, cross-modal retrieval, and unsupervised domain adaptation benchmarks.

ABSTRACT

Extended Footage for EPIC-KITCHENS dataset, to 100 hours of footage. For automatic annotations, see separate dataset at: https://doi.org/10.5523/bris.3l8eci2oqgst92n14w2yqi5ytu 10/09/2020 **N.b. please also see ERRATUM published at https://github.com/epic-kitchens/epic-kitchens-100-annotations/blob/master/README.md#erratum**

Motivation & Objective

  • Extend EPIC-KITCHENS to 100 hours of unscripted egocentric video across 45 environments.
  • Provide a denser, more complete annotation pipeline for fine-grained actions.
  • Enable and define multiple benchmarks (action recognition, action detection, anticipation, retrieval, unsupervised domain adaptation) with baselines and metrics.
  • Examine generalisation over time and domain gaps (test of time) and scalability with additional data.

Proposed method

  • Introduce a scalable narration-based annotation pipeline (pause-and-talk) to densify action annotations.
  • Re-parse and refine verb/noun taxonomies and cluster them into minimal overlapping classes.
  • Annotate temporal bounds of action segments via a larger, crowd-sourced annotator setup and quality controls.
  • Enrich annotations with automatic spatial priors using Mask R-CNN and hand-object interaction detectors.
  • Define six challenges with baselines and metrics, and publish scripts and models for reproducibility.
  • Perform data splits (Train/Val/Test) with unseen participants and tail classes to stress generalisation.
Figure 1: Left: Frames from EPIC-KITCHENS-100 showcasing returning participants with returning or changing kitchens (top) as well as new participants (bottom). Right: Comparisons between recordings from [1] and newly collected videos, with selected frames showcasing the same action. Note object loca
Figure 1: Left: Frames from EPIC-KITCHENS-100 showcasing returning participants with returning or changing kitchens (top) as well as new participants (bottom). Right: Comparisons between recordings from [1] and newly collected videos, with selected frames showcasing the same action. Note object loca

Experimental results

Research questions

  • RQ1How much does denser, multi-level annotation improve action granularity and downstream task performance in egocentric video?
  • RQ2How do models trained on EPIC-KITCHENS-55/earlier data generalise to EPIC-KITCHENS-100, including unseen participants and environments (test of time)?
  • RQ3What is the impact of unsupervised domain adaptation and weak supervision on action recognition and detection in this unscripted dataset?
  • RQ4What are the baseline capabilities for action recognition, detection, anticipation, cross-modal retrieval, and domain adaptation on EPIC-KITCHENS-100?
  • RQ5How does adding new, diverse data affect scalability and generalisation across domain gaps?

Key findings

  • EPIC-KITCHENS-100 comprises 100 hours, 700 videos, and 89,977 fine-grained action segments across 4,053 action classes (verbs, nouns, and actions).
  • The new annotation pipeline yields 54% more actions per minute and 128% more action segments than the previous edition.
  • Verbs (97 classes) and nouns (300 classes) show a long-tail distribution with many new classes present in the newly-collected videos.
  • Unseen-participants and tail-class evaluations demonstrate domain gaps and the need for diverse data and robust models for generalisation.
  • Action detection remains challenging, with mean average precision (mAP) generally low at higher IoU thresholds, highlighting the complexity of overlapping, long, and varied-length actions.
  • The dataset enables six benchmarks, including strong/weak supervision action recognition, action detection, anticipation, retrieval, and unsupervised domain adaptation, with publicly available baselines and evaluation scripts.
Figure 2: Annotation pipeline: (a) narrator, (b) transcriber, (c) temporal segment annotator and (d) dependency parser. Red arrows show AMT crowdsourcing of annotations.
Figure 2: Annotation pipeline: (a) narrator, (b) transcriber, (c) temporal segment annotator and (d) dependency parser. Red arrows show AMT crowdsourcing of annotations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.