Skip to main content
QUICK REVIEW

[論文レビュー] 3D-MuPPET: 3D Multi-Pigeon Pose Estimation and Tracking

Urs Waldmann, Alex Hoi Hang Chan|arXiv (Cornell University)|Aug 29, 2023
Animal Behavior and Welfare Studies被引用数 7
ひとこと要約

3D-MuPPET は、複数のカメラを用いたリアルタイムな 3D マルチピグイン姿勢推定とトラッキングのための新規フレームワークであり、2Dキーポoin検出とトリアングレーション、複数視点間でのアイデンティティトラッキングを統合している。2D では最大 9.45 fps、3D では 1.89 fps の性能を達成し、最先端の手法と同等の精度を発揮するとともに、追加のアノテーションなしに屋外環境へも一般化可能であり、集団行動研究に向けたスケーラブルな 3D 姿勢トラッキングを実現する。

ABSTRACT

Markerless methods for animal posture tracking have been rapidly developing recently, but frameworks and benchmarks for tracking large animal groups in 3D are still lacking. To overcome this gap in the literature, we present 3D-MuPPET, a framework to estimate and track 3D poses of up to 10 pigeons at interactive speed using multiple camera views. We train a pose estimator to infer 2D keypoints and bounding boxes of multiple pigeons, then triangulate the keypoints to 3D. For identity matching of individuals in all views, we first dynamically match 2D detections to global identities in the first frame, then use a 2D tracker to maintain IDs across views in subsequent frames. We achieve comparable accuracy to a state of the art 3D pose estimator in terms of median error and Percentage of Correct Keypoints. Additionally, we benchmark the inference speed of 3D-MuPPET, with up to 9.45 fps in 2D and 1.89 fps in 3D, and perform quantitative tracking evaluation, which yields encouraging results. Finally, we showcase two novel applications for 3D-MuPPET. First, we train a model with data of single pigeons and achieve comparable results in 2D and 3D posture estimation for up to 5 pigeons. Second, we show that 3D-MuPPET also works in outdoors without additional annotations from natural environments. Both use cases simplify the domain shift to new species and environments, largely reducing annotation effort needed for 3D posture tracking. To the best of our knowledge we are the first to present a framework for 2D/3D animal posture and trajectory tracking that works in both indoor and outdoor environments for up to 10 individuals. We hope that the framework can open up new opportunities in studying animal collective behaviour and encourages further developments in 3D multi-animal posture tracking.

研究の動機と目的

  • 屋内および屋外環境の両方において、3D マルチ動物の姿勢推定とトラッキングのためのフレームワークが不足しているという問題に取り組む。
  • 複数視点カメラセットアップを用いて、最大 10 匹のピグインに対してインタラクティブスピードの 3D 姿勢推定と軌道トラッキングを可能にする。
  • ラボデータから野生環境へ、および単一個体から複数個体データへのドメイン一般化を可能にすることで、アノテーション作業を削減する。
  • 3D-POP データセットを用いて、3D マルチ動物の姿勢推定とトラッキングのベンチマークを提供する。
  • 単一動物で訓練された 2D 姿勢推定モデルを、最小限の性能低下で複数個体トラッキングに応用する可能性を実証する。

提案手法

  • キーポイントRCNNに基づく 2D 姿勢推定器が、複数のカメラ視点で複数のピグインのキーポイントとバウンディングボックスを検出する。
  • 複数視点からの 2D キーポイント検出結果を用いてトリアングレーションを適用し、3D 姿勢を再構築する。
  • 最初のフレームにおけるグローバルなアイデンティティに、動的 2D ディテクションマッチングによりアイデンティティを割り当てる。その後のフレームでは 2D トラックヤー(SORT)を用いる。
  • 3D 姿勢推定値の平滑化と時間的整合性の向上のため、カルマンフィルタを用いる。
  • 単一ピグインデータで学習し、再トレーニングなしに複数ピグインおよび屋外環境へ適用可能なドメイン一般化をサポートする。
  • 3D-POP データセットを用いて、中央誤差、正しく検出されたキーポイントの割合(PCK)、およびマルチオブジェクトトラッキング指標を含む定量的指標で評価する。

実験結果

リサーチクエスチョン

  • RQ1単一ピグインで訓練された 2D 姿勢推定モデルが、複数ピグイントラッキングに一般化可能であり、同等の 3D 姿勢精度を達成できるか?
  • RQ2屋内ラボデータで学習した 3D 姿勢推定フレームワークが、追加のアノテーションなしに屋外フィールドデータへも効果的に適用可能か?
  • RQ33D-MuPPET の推論速度と精度は、マルチピグイン状況における最先端の 3D 姿勢推定器と比較してどの程度か?
  • RQ4アイデンティティトラッキングパイプラインは、複数のカメラと時間経過にわたって一貫した ID を維持できるか?
  • RQ5このフレームワークは、リアルタイムかつインタラクティブスピードの 3D マルチ動物の姿勢推定とトラッキングをどの程度サポートできるか?

主な発見

  • 3D-MuPPET は 3D-POP データセットで 2D の中央誤差が 1.85 mm、3D の中央誤差が 2.98 mm を達成し、最先端の手法と同等の精度を発揮する。
  • 2D 推論で 9.45 fps、3D 推論で 1.89 fps の性能を達成し、インタラクティブスピードのトラッキングを実現する。
  • 単一ピグインデータで学習したモデルを用いることで、最大 5 匹のピグインについても 2D および 3D 姿勢推定精度が同等に保たれ、アノテーションの必要性が低減される。
  • 追加のアノテーションなしに屋外環境へも効果的に一般化され、ドメインシフトに対して強固であることが示された。
  • 3D-POP における定量的トラッキング評価では、マルチオブジェクトトラッキングの精度と再現率の両面で良好な結果が得られた。
  • 2D 姿勢推定に KeypointRCNN を用いることで、3D-ViTPose よりも高速な推論が可能となり、速度が求められる応用に適している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。