Skip to main content
QUICK REVIEW

[論文レビュー] A Cross-Modal Distillation Network for Person Re-identification in RGB-Depth.

Frank M. Hafner, Amran Bhuiyan|arXiv (Cornell University)|Oct 27, 2018
Video Surveillance and Tracking Methods参考文献 38被引用数 17
ひとこと要約

本稿では、RGBと深度モダリティ間の構造的特徴を相互に転送するクロスモダリティ蒸留ネットワークを提案し、2段階の最適化により共有埋め込み空間での特徴の整合性を図る。この手法は、計算複雑性を増加させることなく、BIWIおよびRobotPKUデータセットで最先端の性能を達成する。

ABSTRACT

Person re-identification involves the recognition over time of individuals captured using multiple distributed sensors. With the advent of powerful deep learning methods able to learn discriminant representations for visual recognition, cross-modal person re-identification based on different sensor modalities has become viable in many challenging applications in, e.g., autonomous driving, robotics and video surveillance. Although some methods have been proposed for re-identification between infrared and RGB images, few address depth and RGB images. In addition to the challenges for each modality associated with occlusion, clutter, misalignment, and variations in pose and illumination, there is a considerable shift across modalities since data from RGB and depth images are heterogeneous. In this paper, a new cross-modal distillation network is proposed for robust person re-identification between RGB and depth sensors. Using a two-step optimization process, the proposed method transfers supervision between modalities such that similar structural features are extracted from both RGB and depth modalities, yielding a discriminative mapping to a common feature space. Our experiments investigate the influence of the dimensionality of the embedding space, compares transfer learning from depth to RGB and vice versa, and compares against other state-of-the-art cross-modal re-identification methods. Results obtained with BIWI and RobotPKU datasets indicate that the proposed method can successfully transfer descriptive structural features from the depth modality to the RGB modality. It can significantly outperform state-of-the-art conventional methods and deep neural networks for cross-modal sensing between RGB and depth, with no impact on computational complexity.

研究の動機と目的

  • RGBと深度センサー間のクロスモダリティ人物再識別における課題に取り組むこと。これは、モダリティの非均一性と構造的不整合に起因する。
  • 外観、ポーズ、照明、センサー固有のノイズの違いにもかかわらず、RGBと深度モダリティ間での頑健な特徴転送を可能にすること。
  • クロスモダリティ蒸留を通じて共有で判別性の高い特徴空間を学習し、人物再識別の精度を向上させること。
  • 埋め込み次元と転送方向(深度→RGB 対 RGB→深度)がクロスモダリティ学習に与える影響を評価すること。
  • RGBおよび深度データを用いた既存の最先端手法を上回るクロスモダリティ再識別性能を達成すること。

提案手法

  • RGBと深度モダリティ間の監視を転送するために2段階の最適化プロセスを用い、構造的特徴の整合性を図る。
  • 知識蒸留を用いて、一方のモダリティからもう一方のモダリティへ判別性の高い表現を転送し、共有特徴空間における一貫性を促進する。
  • 両モダリティの特徴が共通の判別性の高い表現にマップされる共有埋め込み空間を学習する。
  • RGBと深度モダリティの対応する特徴間の乖離を最小化するようにネットワークを訓練し、クロスモダリティ一般化を強化する。
  • 性能向上を実現しながらも、計算複雑性を維持するように設計されている。
  • 双方向転送をサポートしており、深度→RGBおよびRGB→深度の特徴転送の評価が可能である。

実験結果

リサーチクエスチョン

  • RQ1埋め込み空間の次元がクロスモダリティ人物再識別性能にどのように影響するか?
  • RQ2特徴転送の方向性(深度→RGB 対 RGB→深度)のうち、どちらがより優れた性能を示すか?
  • RQ3提案された蒸留ネットワークは、既存の最先端手法を上回る性能を発揮できるか?
  • RQ4この手法は、異種データ間でモダリティシフトをどの程度軽減し、特徴の整合性を向上させられるか?
  • RQ5優れた精度を達成する一方で、計算複雑性を低く維持できるか?

主な発見

  • 提案手法は、BIWIおよびRobotPKUデータセットにおいて、従来の従来的および深層学習ベースのクロスモダリティ再識別手法を顕著に上回る性能を達成する。
  • 深度モダリティからRGBモダリティへの構造的特徴の転送が、再識別精度の向上に寄与し、蒸留プロセスの有効性を示している。
  • 計算複雑性を増加させることなく最先端の性能を達成し、効率性を維持している。
  • 2段階の最適化プロセスは、モダリティシフトを効果的に低減し、共有埋め込み空間における特徴の整合性を向上させた。
  • アブレーションスタディにより、埋め込み次元の選択が性能に影響することが確認され、最適な設定が最良の結果をもたらした。
  • 双方向転送の実験から、深度→RGB転送の方がRGB→深度転送よりも一般的に優れた性能を示しており、深度特徴の判別力の高さが顕著に表れている。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。