Skip to main content
QUICK REVIEW

[論文レビュー] Deep Gated Multi-modal Learning: In-hand Object Pose Changes Estimation using Tactile and Image Data

Tomoki Anzai, Kuniyuki Takahashi|arXiv (Cornell University)|Sep 27, 2019
Robot Manipulation and Learning参考文献 38被引用数 6
ひとこと要約

本稿では、RGB画像とGelSightタクトイルデータを用いて、各モダリティの信頼性を自己決定的に特定することで、インハンドオブジェクトポーズの変化を動的に推定するエンドツーエンドのディープラーニングフレームワークであるDeep Gated Multi-modal Learning (DGML)を提案する。本手法は、3Dオブジェクトモデルを必要とせず、遮蔽、ノイズ、センサ障害の下でもロバストなポーズ推定を達成し、学習可能なゲート機構による適応的モダリティ重み付けにより、未知のオブジェクトへの一般化性能を向上させる。

ABSTRACT

For in-hand manipulation, estimation of the object pose inside the hand is one of the important functions to manipulate objects to the target pose. Since in-hand manipulation tends to cause occlusions by the hand or the object itself, image information only is not sufficient for in-hand object pose estimation. Multiple modalities can be used in this case, the advantage is that other modalities can compensate for occlusion, noise, and sensor malfunctions. Even though deciding the utilization rate of a modality (referred to as reliability value) corresponding to the situations is important, the manual design of such models is difficult, especially for various situations. In this paper, we propose deep gated multi-modal learning, which self-determines the reliability value of each modality through end-to-end deep learning. For the experiments, an RGB camera and a GelSight tactile sensor were attached to the parallel gripper of the Sawyer robot, and the object pose changes were estimated during grasping. A total of 15 objects were used in the experiments. In the proposed model, the reliability values of the modalities were determined according to the noise level and failure of each modality, and it was confirmed that the pose change was estimated even for unknown objects.

研究の動機と目的

  • ロボット操作における遮蔽、ノイズ、センサ障害下でのインハンドオブジェクトポーズ推定の課題に対処すること。
  • マルチモダリティセンサ統合のための手動で設計された信頼性モデルの必要性を排除すること。
  • 事前3Dモデルの仮定なしに、既知および未知のオブジェクトの両方に対して正確なポーズ変化推定を可能にすること。
  • 入力品質および環境条件に基づき、モダリティの寄与度を自動的に調整する動的でエンドツーエンドの学習フレームワークの開発

提案手法

  • 本手法は、Sawyerロボットグリッパーに取り付けられたGelSightセンサからの順序付きRGB画像とタクトイルデータを処理するディープニューラルネットワークを用いる。
  • 画像モダリティとタクトイルモダリティの信頼性値を出力する学習可能なゲート機構(αおよびβ)を導入し、動的モダリティ重み付けを可能にする。
  • ネットワークはエンドツーエンドに訓練され、初期グリップからの相対的な時間的変化(並進および回転)を予測する。
  • ゲート値はトレーニング中に学習され、ノイズや遮蔽などの入力品質に基づいた各モダリティの信頼性に関するネットワークの自信を反映する。
  • アーキテクチャは、ゲート値が各モダリティの寄与度を決定する重み付き和により、モダリティ固有の特徴を統合する。
  • 本手法は3Dオブジェクトモデルを必要とせず、未観測のオブジェクトへの一般化を可能にする。

実験結果

リサーチクエスチョン

  • RQ1ディープラーニングモデルは、インハンドオブジェクトポーズ推定の過程で、画像およびタクトイルモダリティの信頼性を動的かつ自動的に特定できるか?
  • RQ2本手法は、ノイズ、センサ障害、遮蔽のレベルが異なる条件下でどのように性能を発揮するか?
  • RQ33Dモデル情報が事前に与えられていない状況下で、モデルは未知のオブジェクトにどの程度一般化できるか?
  • RQ4学習された信頼性値は、実際のセンサ品質および形状・サイズなどのオブジェクト特性とどの程度相関しているか?

主な発見

  • ノイズによってタクトイル入力が劣化した場合、ネットワークが画像の信頼性(α)を高めるよう学習したことが確認され、劣化したモダリティ品質の正しく認識されていることが示された。
  • 多くのオブジェクトにおいて、タクトイルデータが画像データよりも信頼性が高かった。特に形状が複雑なオブジェクトや表面特徴が豊富なオブジェクトでは、α値が0.5未満となるケースが多かった。
  • 表面に不規則性が強いオブジェクト(例:テープ、レンチ、コップ)では、画像の信頼性が低く(α < 0.2)、タクトイル特徴への依存度が高くなる傾向が示された。
  • 自己遮蔽やカメラ視界の一部喪失を引き起こす大型オブジェクトでは、画像の信頼性が低下し、ネットワークがタクトイル入力を優先する傾向が見られた。
  • 未知のオブジェクトに対してもロバストなポーズ推定が達成され、3Dモデル依存性のない一般化が確認された。
  • 信頼性ゲート値は、ネットワークの挙動に関する解釈可能なインサイトを提供し、センサ劣化やオブジェクト固有の課題に対して一貫した反応を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。