Skip to main content
QUICK REVIEW

[論文レビュー] AMIL: Adversarial Multi Instance Learning for Human Pose Estimation

Pourya Shamsolmoali, Masoumeh Zareapoor|arXiv (Cornell University)|Mar 18, 2020
Human Pose and Action Recognition参考文献 65被引用数 5
ひとこと要約

本稿では、生成的敵対ネットワーク(GAN)を用い、人体の構造的事前知識を統合することで、ヒューマンポーズ推定を向上させる、新しい敵対的マルチインスタンス学習フレームワークであるAMILを提案する。2つの残差MILモデル(1つの生成器と1つの識別器)を用いる。識別器は本物のヒートマップと生成されたヒートマップを区別し、敵対的学習によって生成器がより現実的で構造的に妥当なポーズを生成するよう誘導される。この手法により、ベンチマークデータセット上での精度が顕著に向上する。

ABSTRACT

Human pose estimation has an important impact on a wide range of applications from human-computer interface to surveillance and content-based video retrieval. For human pose estimation, joint obstructions and overlapping upon human bodies result in departed pose estimation. To address these problems, by integrating priors of the structure of human bodies, we present a novel structure-aware network to discreetly consider such priors during the training of the network. Typically, learning such constraints is a challenging task. Instead, we propose generative adversarial networks as our learning model in which we design two residual multiple instance learning (MIL) models with the identical architecture, one is used as the generator and the other one is used as the discriminator. The discriminator task is to distinguish the actual poses from the fake ones. If the pose generator generates the results that the discriminator is not able to distinguish from the real ones, the model has successfully learnt the priors. In the proposed model, the discriminator differentiates the ground-truth heatmaps from the generated ones, and later the adversarial loss back-propagates to the generator. Such procedure assists the generator to learn reasonable body configurations and is proved to be advantageous to improve the pose estimation accuracy. Meanwhile, we propose a novel function for MIL. It is an adjustable structure for both instance selection and modeling to appropriately pass the information between instances in a single bag. In the proposed residual MIL neural network, the pooling action adequately updates the instance contribution to its bag. The proposed adversarial residual multi-instance neural network that is based on pooling has been validated on two datasets for the human pose estimation task and successfully outperforms the other state-of-arts models.

研究の動機と目的

  • 混雑または複雑なシーンにおける関節の隠蔽や身体部位の重なりによって生じるポーズ推定の不正確さという課題に対処する。
  • 深層学習モデルに人体の構造的事前知識を統合し、ポーズ推定のロバスト性を向上させる。
  • バッチ(画像)内におけるインスタンス(キーポイント)間の関係を効果的にモデル化できる、新しい残差MILアーキテクチャを開発する。
  • 敵対的学習を活用して、生成器がより現実的で解剖学的に妥当なヒートマップを生成するよう誘導する。
  • 標準ベンチマーク上での最先端手法を上回る一般化性能と精度を向上させる。

提案手法

  • 生成器と識別器の両方に共通のアーキテクチャを持つ残差マルチインスタンス学習(MIL)ネットワークを、GANフレームワーク内で提案する。
  • 識別器を用いて、本物の正解ヒートマップと生成器が生成した偽のヒートマップを区別する。
  • 生成器を敵対的に学習させ、本物のヒートマップと区別がつかないようなヒートマップを生成するよう学習させ、結果として構造的事前知識を組み込む。
  • MIL層に、インスタンスの寄与度をバッチ内ですべてのインスタンスに適応的に更新する新しいプーリング機構を設計する。これにより、情報伝達が向上する。
  • 敵対的損失と標準的なポーズ推定損失を組み合わせ、リアルさと正確さの両方を同時に最適化する。
  • 画像入力をエンドツーエンドで処理し、人体の14キーポイント位置のヒートマップを予測する。

実験結果

リサーチクエスチョン

  • RQ1識別器による敵対的学習は、予測されたヒューマンポーズのヒートマップの現実性と構造的一致性を向上させるか?
  • RQ2提案された残差MIL層は、1枚の画像(バッチ)内でのインスタンス間関係をモデル化する上で効果的か?
  • RQ3敵対的学習による構造的事前知識の統合は、隠蔽や重なりがある状況でもポーズ推定精度に測定可能な向上をもたらすか?
  • RQ4提案されたAMILモデルは、標準的なヒューマンポーズ推定ベンチマークにおいて、最先端手法と比較してどうか?
  • RQ5提案されたMILプーリング機構は、キーポイントインスタンスの重みを適応的に設定することで、特徴表現を向上させるか?

主な発見

  • 提案されたAMILモデルは、2つの標準的なヒューマンポーズ推定データセットで最先端の性能を達成し、既存手法を上回る。
  • 敵対的学習により、特に隠蔽や重なりのある状況下でも、予測されたキーポイントヒートマップの構造的一致性が顕著に向上する。
  • 適応的プーリングを備えた残差MIL層は、バッチ内でのインスタンス寄与度を動的に調整することで、特徴表現を強化する。
  • 識別器が本物のヒートマップと生成されたヒートマップを区別できることから、生成器が意味のある構造的事前知識を学習していることが確認できる。
  • 関節の隠蔽や身体部位の重なりに対してもロバストであることが示され、より正確で一貫性のあるポーズ予測が得られる。
  • 定量的評価では、MPIIおよびCOCOデータセットにおいて、ベースラインおよびSOTAモデルと比較して、mAP(平均平均精度)とキーポイント精度が向上している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。