Skip to main content
QUICK REVIEW

[Paper Review] Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events

Weiyao Lin, Huabin Liu|arXiv (Cornell University)|May 9, 2020
Human Pose and Action Recognition44 references67 citations
TL;DR

HiEve introduces a large-scale, hierarchical dataset for human-centric video analysis in crowded and complex events, with extensive pose, tracking, and action annotations, plus cross-annotation baselines and an online evaluation server.

ABSTRACT

Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A complex event relates to dense crowds, anomalous individuals, or collective behaviors. However, limited by the scale and coverage of existing video datasets, few human analysis approaches have reported their performances on such complex events. To this end, we present a new large-scale dataset with comprehensive annotations, named Human-in-Events or HiEve (Human-centric video analysis in complex Events), for the understanding of human motions, poses, and actions in a variety of realistic events, especially in crowd & complex events. It contains a record number of poses (>1M), the largest number of action instances (>56k) under complex events, as well as one of the largest numbers of trajectories lasting for longer time (with an average trajectory length of >480 frames). Based on its diverse annotation, we present two simple baselines for action recognition and pose estimation, respectively. They leverage cross-label information during training to enhance the feature learning in corresponding visual tasks. Experiments show that they could boost the performance of existing action recognition and pose estimation pipelines. More importantly, they prove the widely ranged annotations in HiEve can improve various video tasks. Furthermore, we conduct extensive experiments to benchmark recent video analysis approaches together with our baseline methods, demonstrating HiEve is a challenging dataset for human-centric video analysis. We expect that the dataset will advance the development of cutting-edge techniques in human-centric analysis and the understanding of complex events. The dataset is available at http://humaninevents.org

Motivation & Objective

  • Build a large-scale dataset of real-world complex events focusing on human motions, poses, and actions in crowded scenes.
  • Provide comprehensive annotations (pose, track, action) to enable multiple video understanding tasks.
  • Demonstrate cross-annotation benefits by designing pose-aware action recognition and action-guided pose estimation baselines.
  • Evaluate state-of-the-art methods on HiEve to establish its challenge level and baselines’ impact.

Proposed method

  • Curate 12 real-world scenes with diverse complex events and collect 32 video sequences totaling 49,820 frames.
  • Annotate each person with 14 keypoints (nose, chest, shoulders, elbows, wrists, hips, knees, ankles) across frames, including invisible keypoints when needed.
  • Annotate 14 action categories every 20 frames for all individuals, handling group actions by annotating all participants.
  • Provide dense annotations enabling tracking, pose estimation, and action recognition, plus long trajectories (average length >480 frames).
  • Develop two enhanced baselines that leverage cross-annotation information: (i) pose-aware action recognition, integrating pose features into RGB-based action models, and (ii) action-guided pose estimation, using action priors to refine poses.
  • Introduce evaluation metrics and an online server (HiEve evaluation server) for held-out test videos.

Experimental results

Research questions

  • RQ1How does HiEve’s scale and annotation diversity support robust evaluation and development of real-world human-centric video analysis methods?
  • RQ2Can cross-annotation information (pose, track, action) improve performance on action recognition and pose estimation in crowded, complex events?
  • RQ3How do state-of-the-art methods perform on HiEve’s complex scenes compared to existing benchmarks?
  • RQ4What are the challenges of long-term re-identification and crowded scene understanding as captured by HiEve?

Key findings

  • HiEve contains 49,820 frames, 1,099,357 poses, 56,643 action instances, and 2,687 long trajectories with an average length of 485 frames.
  • HiEve captures longer, more crowded scenes than several prior MOT and pose datasets, demonstrating increased difficulty for tracking and pose estimation in complex events.
  • Cross-annotation baselines (pose-aware action recognition and action-guided pose estimation) boost performance of existing pipelines on HiEve.
  • HiEve’s diverse annotations and challenging scenarios improve evaluation of current video analysis approaches, and the authors provide an online evaluation server to enable scalable benchmarking.
  • HiEve is positioned as a challenging benchmark for advancing human-centric video analysis in realistic crowded environments.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.