Skip to main content
QUICK REVIEW

[论文解读] Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events

Weiyao Lin, Huabin Liu|arXiv (Cornell University)|May 9, 2020
Human Pose and Action Recognition参考文献 44被引用 67
一句话总结

HiEve 引入一个面向人类中心的视频分析的大型分层数据集,覆盖拥挤且复杂事件中的姿态、跟踪和动作注释,并提供跨注释基线和在线评估服务器。

ABSTRACT

Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A complex event relates to dense crowds, anomalous individuals, or collective behaviors. However, limited by the scale and coverage of existing video datasets, few human analysis approaches have reported their performances on such complex events. To this end, we present a new large-scale dataset with comprehensive annotations, named Human-in-Events or HiEve (Human-centric video analysis in complex Events), for the understanding of human motions, poses, and actions in a variety of realistic events, especially in crowd & complex events. It contains a record number of poses (>1M), the largest number of action instances (>56k) under complex events, as well as one of the largest numbers of trajectories lasting for longer time (with an average trajectory length of >480 frames). Based on its diverse annotation, we present two simple baselines for action recognition and pose estimation, respectively. They leverage cross-label information during training to enhance the feature learning in corresponding visual tasks. Experiments show that they could boost the performance of existing action recognition and pose estimation pipelines. More importantly, they prove the widely ranged annotations in HiEve can improve various video tasks. Furthermore, we conduct extensive experiments to benchmark recent video analysis approaches together with our baseline methods, demonstrating HiEve is a challenging dataset for human-centric video analysis. We expect that the dataset will advance the development of cutting-edge techniques in human-centric analysis and the understanding of complex events. The dataset is available at http://humaninevents.org

研究动机与目标

  • 构建一个聚焦于拥挤场景中人类动作、姿态和动作的现实世界复杂事件的大规模数据集。
  • 提供全面的注释(姿态、跟踪、动作),以支持多种视频理解任务。
  • 通过设计姿态感知的动作识别和动作引导的姿态估计基线来展示跨注释的优势。
  • 在 HiEve 上评估最先进的方法,以确立其挑战难度和基线影响。

提出的方法

  • 整理 12 个具有多样化复杂事件的现实世界场景,并收集 32 条视频序列,总计 49,820 帧。
  • 在每一帧中为每个人标注 14 个关键点(鼻子、胸部、肩膀、肘部、手腕、髋部、膝盖、踝部),在需要时包含不可见的关键点。
  • 每 20 帧对所有个体标注 14 个动作类别,对于群体动作通过标注所有参与者来处理。
  • 提供密集注释,支持跟踪、姿态估计和动作识别,以及较长的轨迹(平均长度 >480 帧)。
  • 开发两个利用跨注释信息的增强型基线:(i) 姿态感知的动作识别,将姿态特征并入基于 RGB 的动作模型,(ii) 动作引导的姿态估计,利用动作先验来细化姿态。
  • 为保留测试视频引入评估指标和一个在线服务器(HiEve evaluation server)。

实验结果

研究问题

  • RQ1HiEve 的规模和注释多样性如何支持对现实世界人类中心视频分析方法的稳健评估与开发?
  • RQ2跨注释信息(姿态、跟踪、动作)是否能提升在拥挤、复杂事件中的动作识别和姿态估计性能?
  • RQ3最先进方法在 HiEve 的复杂场景中的表现与现有基准相比如何?
  • RQ4如 HiEve 所示,长期重新识别和拥挤场景理解面临的挑战有哪些?

主要发现

  • HiEve 包含 49,820 帧、1,099,357 个姿态、56,643 个动作实例,以及 2,687 条平均长度为 485 帧的长轨迹。
  • HiEve 捕捉的场景比若干早期 MOT 与姿态数据集更长且更拥挤,这表明在复杂事件中跟踪和姿态估计的难度更高。
  • 跨注释基线(姿态感知的动作识别和动作引导的姿态估计)提升了现有管线在 HiEve 上的性能。
  • HiEve 多样化的注释和具有挑战性的场景提升了对当前视频分析方法的评估,作者提供在线评估服务器以实现可扩展的基准测试。
  • HiEve 被定位为在现实拥挤环境中推进面向人类的视频分析的具有挑战性的基准。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。