Skip to main content
QUICK REVIEW

[論文レビュー] LISA: Learning Implicit Shape and Appearance of Hands

Enric Corona, Tomáš Hodaň|arXiv (Cornell University)|Apr 4, 2022
Human Pose and Action Recognition被引用数 4
ひとこと要約

LISA は、粗い 3D 手関節ポーズを伴うマルチビュー RGB 動画から、正確な 3D 形状、外観、ポーズを同時に学習するニューラルインクリメントモデルである。各ボーンごとの符号付き距離関数と色予測を、学習されたスキンニング重みを介して組み合わせることで、LISA は画像および点群からの最先端の再構成品質を達成し、形状、色、ポーズパラメータの分離制御を可能にする。

ABSTRACT

This paper proposes a do-it-all neural model of human hands, named LISA. The model can capture accurate hand shape and appearance, generalize to arbitrary hand subjects, provide dense surface correspondences, be reconstructed from images in the wild and easily animated. We train LISA by minimizing the shape and appearance losses on a large set of multi-view RGB image sequences annotated with coarse 3D poses of the hand skeleton. For a 3D point in the hand local coordinate, our model predicts the color and the signed distance with respect to each hand bone independently, and then combines the per-bone predictions using predicted skinning weights. The shape, color and pose representations are disentangled by design, allowing to estimate or animate only selected parameters. We experimentally demonstrate that LISA can accurately reconstruct a dynamic hand from monocular or multi-view sequences, achieving a noticeably higher quality of reconstructed hand shapes compared to baseline approaches. Project page: https://www.iri.upc.edu/people/ecorona/lisa/.

研究の動機と目的

  • 多様な被験者に対して詳細な手の形状と外観を捉えるニューラルインクリメントモデルの開発。
  • 部分的遮蔽がある状況下でも、単眼またはマルチビュー画像からの正確な 3D 手再構成の実現。
  • より良いアニメーションと制御を実現するため、明示的に予測されたスキンニング重みによる密な表面対応の提供。
  • 形状、色、ポーズ表現の分離を実現し、細かく制御可能で一般化性の高いモデルの構築。
  • 合成データおよび実世界データの両方で、既存のパラメトリックメッシュモデルやインクリメントモデルを上回る再構成精度の達成。

提案手法

  • エンドツーエンドの学習を可能にするために、粗い 3D 手関節スケルトンアノテーションを備えたマルチビュー RGB 動画データセットを活用。
  • 手を剛体ボーンの集合として表現し、各ボーンがその局所座標系内で SDF と色を独立して予測するニューラルネットワークを備える。
  • 予測されたスキンニング重みを用いて、各ボーンごとの SDF と色予測を統合されたインクリメント表現に統合。
  • 微分可能レンダリングパイプラインを用いて、形状および外観の損失を学習中に最適化。
  • 2段階の最適化を採用:まず 3D ポーズを精緻化し、その後形状、ポーズ、色パラメータを同時に最適化。
  • 可視化およびレンダリングのため、128³ 解像度の高解像度メッシュを再構成。
Figure 2 : Training and architecture of the LISA hand model. Left: LISA is trained by minimizing shape and appearance losses from a dataset of multi-view RGB image sequences. The sequences are assumed annotated with coarse 3D poses of the hand skeleton that are refined during training. The training
Figure 2 : Training and architecture of the LISA hand model. Left: LISA is trained by minimizing shape and appearance losses from a dataset of multi-view RGB image sequences. The sequences are assumed annotated with coarse 3D poses of the hand skeleton that are refined during training. The training

実験結果

リサーチクエスチョン

  • RQ1ニューラルインクリメント表現は、分離制御が可能な複雑でアーチレートな手の幾何学的形状と外観を効果的にモデル化できるか?
  • RQ2このようなモデルは、制約のない画像から、未観測の手の被験者および任意のポーズに一般化できるか?
  • RQ3MANO や NARF や NASA のようなパラメトリックメッシュモデルやインクリメントモデルと比較して、優れた再構成品質を達成できるか?
  • RQ4学習されたスキンニング重みを用いたボーンごとの予測は、表面対応性と再構成忠実度を向上させるか?
  • RQ5モデルは、単眼画像や遮蔽のある実環境シナリオからの手再構成をどの程度正確に実現できるか?

主な発見

  • LISA は、点群からの 3D 手再構成において最先端のパフォーマンスを達成し、真値スキャンへのフィッティングで 1mm 未塔の誤差を達成。
  • DeepHandMesh データセットにおいて、LISA-im は画像のみのベースラインをすべて上回り、1 ビューで PSNR 27.45、2 ビューで 28.40、4 ビューで 28.27 を記録。
  • InterHand2.6M の画像からの色再構成において、LISA-full は PSNR 28.27、SSIM 0.96、LPIPS 0.05 を達成し、知覚的品質において NASA や NARF を上回る。
  • LISA は、部分的遮蔽や複雑な照明条件を効果的に処理し、高精細な実環境再構成を実現。
  • 1 枚のビューあたり単一の P100 GPU で約 1 分で高品質な新規ビューレンダリングを実現。
  • LISA の分離表現により、形状、色、ポーズを独立して操作可能となり、正確な制御とアニメーションが可能になる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。