Skip to main content
QUICK REVIEW

[論文レビュー] A Novel Approach for Robust Multi Human Action Recognition and Summarization based on 3D Convolutional Neural Networks

Noor Almaadeed, Omar Elharrouss|arXiv (Cornell University)|Jul 25, 2019
Human Pose and Action Recognition被引用数 8
ひとこと要約

本論文は、監視映像およびウェブ動画におけるロバストな多人数行動認識および動画要約のための3次元畳み込みニューラルネットワーク(3D CNN)ベースのフレームワークを提案する。本手法はシーンから個々の行動シーケンスを抽出し、3D CNNを用いて行動検出を実行し、行動中心の要約を生成する。UCF101およびYouTubeデータセットにおいて、前処理を一切行わずに最先端の性能を達成している。

ABSTRACT

Human actions in videos are 3D signals. However, there are a few methods available for multiple human action recognition. For long videos, it's difficult to search within a video for a specific action and/or person. For that, this paper proposes a new technic for multiple human action recognition and summarization for surveillance videos. The proposed approach proposes a new representation of the data by extracting the sequence of each person from the scene. This is followed by an analysis of each sequence to detect and recognize the corresponding actions using 3D convolutional neural networks (3DCNNs). Action-based video summarization is performed by saving each person's action at each time of the video. Results of this work revealed that the proposed method provides accurate multi human action recognition that easily used for summarization of any action. Further, for other videos that can be collected from the internet, which are complex and not built for surveillance applications, the proposed model was evaluated on some datasets like UCF101 and YouTube without any preprocessing. For this category of videos, the summarization is performed on the video sequences by summarizing the actions in each subsequence. The results obtained demonstrate its efficiency compared to state-of-the-art methods.

研究の動機と目的

  • 長時間の監視映像における複数人の行動を特定・要約する課題に対処すること。
  • 長大な動画コンテンツ内での特定の行動や人物の効率的検索・取得を可能にする手法を開発すること。
  • UCF101およびYouTubeのようなインターネットからの複雑で非構造的な動画に対しても、行動認識モデルの適用範囲を拡張すること。
  • 各個人の重要な出来事を保持する、コンパクトで行動ベースの動画要約を生成すること。

提案手法

  • 本手法は、人物固有の空間時間的トラッキングを用いて、動画フレームから個々の人物の行動シーケンスを抽出する。
  • 各抽出シーケンスから空間時間的特徴を学習するため、3次元畳み込みニューラルネットワーク(3DCNN)を適用する。
  • 行動ベースの動画要約は、各個人ごとに関連するタイムスタンプでの重要な行動クリップを選択・保持することで実行する。
  • 本アプローチは、UCF101およびYouTubeからの前処理を施さない動画で評価され、データ固有の前処理に依存しないロバスト性を示している。
  • 本フレームワークは、監視映像の要約と、制約のない環境における一般的な動画理解の両方をサポートする。

実験結果

リサーチクエスチョン

  • RQ13D CNNベースの手法は、事前にセグメンテーションされたデータに依存せずに、長時間の監視映像における多人数行動認識を正確に達成できるか?
  • RQ2提案手法は、非構造的なウェブ動画における複数人の行動要約において、どの程度効果的か?
  • RQ3UCF101およびYouTubeのような複雑で制約のない動画データセットに対して、前処理なしでどの程度一般化性能を示すか?
  • RQ4行動ベースの要約技術は、特定の行動や人物のための効率的な動画検索・取得を可能にするか?

主な発見

  • 提案手法は、動画の前処理を一切要せず、UCF101およびYouTubeデータセットにおける多人数行動認識で最先端の性能を達成している。
  • モデルは長時間の監視映像における行動認識に成功し、各個人に対して正確な行動中心の要約を生成している。
  • 本手法は、監視用途に本来設計されていない動画環境を含む、複雑で制約のない動画環境においてもロバストであることを示している。
  • 行動ベースの要約により、各人物ごとに最も関連性の高い行動セグメントのみを保持することで、効率的な動画ナビゲーションが可能になった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。