Skip to main content
QUICK REVIEW

[論文レビュー] Grasping Field: Learning Implicit Representations for Human Grasps

Korrawe Karunratanakul, Jinlong Yang|arXiv (Cornell University)|Aug 10, 2020
Robot Manipulation and Learning参考文献 89被引用数 12
ひとこと要約

本稿では、3次元点を手および物体の表面からの符号付き距離にマッピングすることで、3次元手-物体相互作用をモデル化する深層学習ベースの暗黙的表現であるGrasping Fieldを紹介する。この手法により、未学習の3次元物体点群から高品質で物理的に妥当な人間の grasp を生成でき、単一のRGB画像からの3次元手-物体再構成においても、先行する最先端手法を上回る接触の妥当性と再構成品質を達成する。

ABSTRACT

Robotic grasping of house-hold objects has made remarkable progress in recent years. Yet, human grasps are still difficult to synthesize realistically. There are several key reasons: (1) the human hand has many degrees of freedom (more than robotic manipulators); (2) the synthesized hand should conform to the surface of the object; and (3) it should interact with the object in a semantically and physically plausible manner. To make progress in this direction, we draw inspiration from the recent progress on learning-based implicit representations for 3D object reconstruction. Specifically, we propose an expressive representation for human grasp modelling that is efficient and easy to integrate with deep neural networks. Our insight is that every point in a three-dimensional space can be characterized by the signed distances to the surface of the hand and the object, respectively. Consequently, the hand, the object, and the contact area can be represented by implicit surfaces in a common space, in which the proximity between the hand and the object can be modelled explicitly. We name this 3D to 2D mapping as Grasping Field, parameterize it with a deep neural network, and learn it from data. We demonstrate that the proposed grasping field is an effective and expressive representation for human grasp generation. Specifically, our generative model is able to synthesize high-quality human grasps, given only on a 3D object point cloud. The extensive experiments demonstrate that our generative model compares favorably with a strong baseline and approaches the level of natural human grasps. Our method improves the physical plausibility of the hand-object contact reconstruction and achieves comparable performance for 3D hand reconstruction compared to state-of-the-art methods.

研究の動機と目的

  • 複雑な人間の手-物体相互作用をモデル化する表現力と効率性を備えた表現を開発すること。
  • 未学習の3次元物体点群から、現実的で物理的に妥当な人間の grasp を生成できること。
  • 重ね合わせの禁止などの物理的制約を強制することで、単一のRGB画像からの3次元手および物体の再構成を改善すること。
  • メッシュベースの表現の限界、例えば固定された接触領域や genus の制約を解消すること。
  • 暗黙的で微分可能な接触モデリングを用いて、手-物体相互作用のエンドツーエンド学習を促進すること。

提案手法

  • Grasping Fieldは、3次元点を手および物体の表面からの符号付き距離にマッピングする3次元から2次元への写像として定義される。
  • この表現は、実際の人間の grasp データ上でエンドツーエンドに訓練された深層ニューラルネットワークによってパラメータ化される。
  • 両方の符号付き距離がほぼゼロである領域が接触領域として暗黙的に定義され、自然な手-物体接触モデリングが可能になる。
  • 2デコーダーのネットワークアーキテクチャが採用され、手および物体の符号付き距離を同時に予測する。物理的妥当性を強制する損失項が含まれる。
  • トレーニング中に重なりを最小限に抑え、接触の現実性を向上させるために、接触および重ね合わせ損失が組み込まれている。
  • RGBからの再構成のため、ネットワークはグリッピングフィールドを暗黙的形状表現として用いて、手および物体の幾何を同時に予測するように変更されている。

実験結果

リサーチクエスチョン

  • RQ1学習された暗黙的表現は、3次元物体との複雑で高自由度の手の相互作用を効果的にモデル化できるか?
  • RQ2明示的なメッシュ化やヒューリスティックな接触領域を用いずに、合成された人間の grasp における物理的妥当性と意味的現実性をどのように強制できるか?
  • RQ3グリッピングフィールド表現は、メッシュベースのベースラインと比較して、単一のRGB画像からの3次元手および物体の再構成を改善できるか?
  • RQ4グリッピング生成において、未学習の物体にどの程度一般化できるか?
  • RQ5精度および耐性の観点から、グリッピングフィールドにおける暗黙的接触モデリングは、明示的なメッシュベースの接触推論と比較してどのように異なるか?

主な発見

  • 提案された生成モデルは、未学習の物体に対しても、高品質で物理的および意味的に妥当な人間の grasp を3次元物体点群から合成できる。
  • 再構成タスクにおいて、ベースライン [29] と比較して手と物体の重なりを30%削減し、物理的妥当性が著しく向上した。
  • FHBデータセットでは、偽の真値MANOジョイントを基準に評価した際、手関節再構成誤差が2.6 cmにまで低下し、最先端の性能に近づいた。
  • グリッピングフィールド表現により、事前に定義された接触ゾーンやメッシュ解像度の制約が不要な状態で、正確な接触領域の推論が可能になった。
  • 再構成された手のMANOフィッティングの結果、統計的手法による正則化に依存せず、現実的な手の形状が生成された。
  • アブレーションスタディにより、接触および重なり損失が再構成品質を著しく向上させ、特に2デコーダー構造において顕著であった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。