Skip to main content
QUICK REVIEW

[論文レビュー] Pedestrian Attribute Recognition in Video Surveillance Scenarios Based on View-attribute Attention Localization

Wei‐Chen Chen, Xinyi Yu|arXiv (Cornell University)|Jun 11, 2021
Video Surveillance and Tracking Methods参考文献 52被引用数 4
ひとこと要約

本論文では、視点固有のアテンションと領域アテンションを活用して属性局在化と認識精度を向上させる、動画監視における歩行者属性認識のための新規手法View-attribute Attention Localization (VALA) を提案する。視点予測とアテンションメカニズムを統合することで、RAP、RAPv2、PA-100K データセットにおいて最先端の性能を達成し、視点と空間的アテンションの共同監視により、優れた局在化および認識能力を示した。

ABSTRACT

Pedestrian attribute recognition in surveillance scenarios is still a challenging task due to the inaccurate localization of specific attributes. In this paper, we propose a novel view-attribute localization method based on attention (VALA), which utilizes view information to guide the recognition process to focus on specific attributes and attention mechanism to localize specific attribute-corresponding areas. Concretely, view information is leveraged by the view prediction branch to generate four view weights that represent the confidences for attributes from different views. View weights are then delivered back to compose specific view-attributes, which will participate and supervise deep feature extraction. In order to explore the spatial location of a view-attribute, regional attention is introduced to aggregate spatial information and encode inter-channel dependencies of the view feature. Subsequently, a fine attentive attribute-specific region is localized, and regional weights for the view-attribute from different spatial locations are gained by the regional attention. The final view-attribute recognition outcome is obtained by combining the view weights with the regional weights. Experiments on three wide datasets (RAP, RAPv2, and PA-100K) demonstrate the effectiveness of our approach compared with state-of-the-art methods.

研究の動機と目的

  • 監視環境下における歩行者属性認識における属性局在化の不正確さという課題に対処すること。
  • 多視点情報とアテンションメカニズムを活用して関連する身体部位に注目することで、認識性能を向上させること。
  • アテンション監視を介して視点推定と属性局在化を統合的に最適化する統一フレームワークの開発。
  • 属性固有の領域に対してチャネル間の依存関係と空間的アテンションをモデル化することで特徴表現を向上させること。

提案手法

  • 視点予測ブランチが、異なる視点(前、後、左、右)からの属性に関する信頼度を表す4つの視点重みを生成する。
  • 視点重みを特徴マップと統合して、アテンション監視のもとで深層特徴抽出を導く視点-属性表現を形成する。
  • 視点固有の特徴における空間的依存関係とチャネル間関係を符号化するために、領域アテンションを適用する。
  • 領域アテンションを介して空間的に配慮した領域重みを生成し、属性固有の身体部位の局在化を精緻化する。
  • 最終的な属性認識は、視点重みと領域重みを重み付き統合機構で統合することで達成される。
  • モデルは、視点と属性局在化の両方の監視を受けて、3つのベンチマークデータセット(RAP、RAPv2、PA-100K)上でエンドツーエンドで訓練される。

実験結果

リサーチクエスチョン

  • RQ1多視点情報を効果的に活用することで、歩行者認識における属性局在化をどのように改善できるか?
  • RQ2アテンションメカニズムは、監視動画における属性固有の身体部位の局在化をどのように向上させられるか?
  • RQ3視点予測と領域アテンションの共同監視が、認識精度をどの程度向上させるか?
  • RQ4標準ベンチマーク上で、提案手法の視点-属性アテンション機構は、従来のアテンションまたは局在化手法と比べてどのように差をつけるか?

主な発見

  • VALA は RAP、RAPv2、PA-100K データセットにおいて最先端の性能を達成し、既存手法よりも属性認識精度で優れている。
  • 視点重みと領域アテンションの統合により、属性関連身体部位の局在化精度が顕著に向上した。
  • 視点予測ブランチにより、異なる歩行者視点における属性認識に自信に基づいたガイダンスが提供され、モデルのロバスト性が向上した。
  • アブレーションスタディの結果、視点監視と領域アテンションの両方が性能向上に有意に寄与していることが確認された。
  • アテンションに基づく局在化と視点に配慮した特徴学習のおかげで、多様な監視シナリオに強く一般化する能力を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。