[論文レビュー] Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events
HiEve は、人間中心のビデオ分析のための大規模で階層的なデータセットを、混雑した複雑なイベントにおいて導入し、広範な姿勢(ポーズ)、追跡、アクションの注釈に加え、クロス注釈ベースラインとオンライン評価サーバを提供します。
Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A complex event relates to dense crowds, anomalous individuals, or collective behaviors. However, limited by the scale and coverage of existing video datasets, few human analysis approaches have reported their performances on such complex events. To this end, we present a new large-scale dataset with comprehensive annotations, named Human-in-Events or HiEve (Human-centric video analysis in complex Events), for the understanding of human motions, poses, and actions in a variety of realistic events, especially in crowd & complex events. It contains a record number of poses (>1M), the largest number of action instances (>56k) under complex events, as well as one of the largest numbers of trajectories lasting for longer time (with an average trajectory length of >480 frames). Based on its diverse annotation, we present two simple baselines for action recognition and pose estimation, respectively. They leverage cross-label information during training to enhance the feature learning in corresponding visual tasks. Experiments show that they could boost the performance of existing action recognition and pose estimation pipelines. More importantly, they prove the widely ranged annotations in HiEve can improve various video tasks. Furthermore, we conduct extensive experiments to benchmark recent video analysis approaches together with our baseline methods, demonstrating HiEve is a challenging dataset for human-centric video analysis. We expect that the dataset will advance the development of cutting-edge techniques in human-centric analysis and the understanding of complex events. The dataset is available at http://humaninevents.org
研究の動機と目的
- 混雑した場面での人間の動作、姿勢、アクションに焦点を当てた実世界の大規模データセットを構築する。
- 複数のビデオ理解タスクを可能にするため、姿勢(ポーズ)、追跡、アクションの包括的な注釈を提供する。
- ポーズを意識したアクション認識とアクション指向のポーズ推定といった、クロス注釈情報の利点を示すベースラインを設計する。
- HiEve 上の最先端手法を評価し、その難易度とベースラインの影響を明らかにする。
提案手法
- 現実世界の12場景をキュレーションし、様々な複雑なイベントを収集して、合計32個のビデオシーケンス、総 frames は 49,820 フレーム。
- 各人物に 14 のキーポイント(鼻、胸、肩、肘、手首、腰、膝、足首) across frames, including invisible keypoints when needed.
- 全員を対象に20フレームごとに14のアクションカテゴリを注釈し、グループアクションは全参加者を注釈して処理する。
- 追跡、ポーズ推定、アクション認識を可能にする密な注釈を提供し、長い軌跡(平均長 >480 フレーム)も含む。
- クロス注釈情報を活用する2つの強化ベースラインを開発する: (i) ポーズを取り入れたアクション認識、ポーズ特徴をRGBベースのアクションモデルに統合、(ii) アクション指向のポーズ推定、アクション事前情報を用いてポーズを洗練。
- 評価指標と、保留データのテスト動画向けオンラインサーバ(HiEve 評価サーバ)を導入する。
実験結果
リサーチクエスチョン
- RQ1HiEve の規模と注釈の多様性は、現実世界の人間中心ビデオ分析手法の堅牢な評価と開発をどのように支えるか?
- RQ2混雑した複雑なイベントにおけるアクション認識とポーズ推定の性能を、ポーズ・追跡・アクションのクロス注釈情報が改善できるか?
- RQ3既存のベンチマークと比べて、HiEve の複雑なシーンで最先端手法はどの程度の性能を示すか?
- RQ4HiEve が捉える長期再識別と混雑場面の理解にはどのような課題があるか?
主な発見
- HiEve は 49,820 フレーム、1,099,357 ポーズ、56,643 アクションインスタンス、2,687 の長い軌跡(平均長 485 フレーム)を含む。
- HiEve は、過去の MOT やポーズデータセットのいくつかよりも長く、群衆の多いシーンを捉え、複雑なイベントにおける追跡とポーズ推定の難易度の増大を示している。
- クロス注釈ベースライン(ポーズを意識したアクション認識とアクション指向のポーズ推定)は、HiEve で既存パイプラインの性能を向上させる。
- HiEve の多様な注釈と難易度の高いシナリオは、現在の映像解析手法の評価を改善し、著者らはスケーラブルなベンチマーキングを可能にするオンライン評価サーバを提供する。
- HiEve は、現実的で混雑した環境における人間中心のビデオ分析を進展させるための挑戦的なベンチマークとして位置づけられている。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。