[Paper Review] Trespassing the Boundaries: Labeling Temporal Bounds for Object Interactions in Egocentric Video
This paper identifies significant inconsistencies in temporal boundary annotations for egocentric object interactions across three datasets, demonstrating that even small shifts in start/end times reduce recognition accuracy by up to 10% in state-of-the-art models. To address this, the authors propose Rubicon Boundaries—inspired by cognitive psychology—to standardize labeling based on distinct action phases, achieving a 4% accuracy gain on GTEA Gaze+ and improved per-class performance for 55% of action classes.
Manual annotations of temporal bounds for object interactions (i.e. start and end times) are typical training input to recognition, localization and detection algorithms. For three publicly available egocentric datasets, we uncover inconsistencies in ground truth temporal bounds within and across annotators and datasets. We systematically assess the robustness of state-of-the-art approaches to changes in labeled temporal bounds, for object interaction recognition. As boundaries are trespassed, a drop of up to 10% is observed for both Improved Dense Trajectories and Two-Stream Convolutional Neural Network. We demonstrate that such disagreement stems from a limited understanding of the distinct phases of an action, and propose annotating based on the Rubicon Boundaries, inspired by a similarly named cognitive model, for consistent temporal bounds of object interactions. Evaluated on a public dataset, we report a 4% increase in overall accuracy, and an increase in accuracy for 55% of classes when Rubicon Boundaries are used for temporal annotations.
Motivation & Objective
- To investigate the consistency of temporal boundary annotations for object interactions in egocentric video across and within three public datasets.
- To evaluate the robustness of state-of-the-art action recognition models—Improved Dense Trajectories and Two-Stream Convolutional Neural Networks—under perturbations of temporal bounds.
- To propose a cognitive-inspired labeling framework, Rubicon Boundaries, to standardize temporal boundary annotation based on distinct action phases.
- To demonstrate that consistent labeling via Rubicon Boundaries improves recognition accuracy and generalization, even when data augmentation is applied.
Proposed method
- Systematically analyze temporal boundary inconsistencies in three egocentric datasets: GTEA Gaze+, GTEA Gaze+ Extended, and EPFL-CD.
- Apply controlled perturbations to ground truth temporal bounds (within IoU > 0.5) to assess recognition robustness of IDT and 2SCNN models.
- Introduce Rubicon Boundaries, a labeling scheme based on cognitive psychology, defining action phases: pre-actional, actional, and post-actional.
- Re-annotate the GTEA Gaze+ dataset using Rubicon Boundaries, creating two variants: 'act' (actional phase only) and 'full' (entire Rubicon segment).
- Train and evaluate 2SCNN on both conventional and Rubicon-annotated segments under identical conditions to isolate the effect of boundary consistency.
- Compare performance across conventional and Rubicon-annotated data using ground truth and generated segments with varied start/end times.
Experimental results
Research questions
- RQ1How consistent are temporal boundary annotations for object interactions across and within egocentric video datasets?
- RQ2To what extent do variations in labeled temporal bounds affect the performance of state-of-the-art action recognition models?
- RQ3Can a cognitive-inspired labeling framework improve consistency and recognition accuracy in egocentric action recognition?
- RQ4Does using Rubicon Boundaries lead to better generalization and robustness compared to conventional annotations?
Key findings
- Temporal boundary annotations across and within the three egocentric datasets show significant inconsistency, with high variability even for the same action class.
- Recognition accuracy drops by up to 10% for both IDT and 2SCNN when temporal bounds are perturbed, despite IoU > 0.5 with ground truth.
- Using Rubicon Boundaries for annotation increases overall recognition accuracy by 4% on the GTEA Gaze+ dataset compared to conventional annotations.
- For 55% of action classes (23 out of 42), recognition accuracy improves when using the full Rubicon Boundary segment, while only 10 classes show no change.
- The Rubicon-annotated segments demonstrate improved robustness to boundary perturbations, with higher accuracy than conventional-generated segments under the same conditions.
- The improvement in accuracy is attributed solely to better boundary consistency, as data augmentation alone cannot replicate the gains from Rubicon labeling.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.