Skip to main content
QUICK REVIEW

[論文レビュー] ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Tongzhou Mu, Zhan Ling|arXiv (Cornell University)|Jul 30, 2021
3D Shape Modeling and Analysis参考文献 89被引用数 27
ひとこと要約

ManiSkillは、多様な可動部品を持つ一般化可能な3D視覚駆動の操作スキルを対象とした大規模なオープンソースベンチマークを導入し、全物理シミュレータで36k件のデモと4つのタスクを提供します。

ABSTRACT

Object manipulation from 3D visual inputs poses many challenges on building generalizable perception and policy models. However, 3D assets in existing benchmarks mostly lack the diversity of 3D shapes that align with real-world intra-class complexity in topology and geometry. Here we propose SAPIEN Manipulation Skill Benchmark (ManiSkill) to benchmark manipulation skills over diverse objects in a full-physics simulator. 3D assets in ManiSkill include large intra-class topological and geometric variations. Tasks are carefully chosen to cover distinct types of manipulation challenges. Latest progress in 3D vision also makes us believe that we should customize the benchmark so that the challenge is inviting to researchers working on 3D deep learning. To this end, we simulate a moving panoramic camera that returns ego-centric point clouds or RGB-D images. In addition, we would like ManiSkill to serve a broad set of researchers interested in manipulation research. Besides supporting the learning of policies from interactions, we also support learning-from-demonstrations (LfD) methods, by providing a large number of high-quality demonstrations (~36,000 successful trajectories, ~1.5M point cloud/RGB-D frames in total). We provide baselines using 3D deep learning and LfD algorithms. All code of our benchmark (simulator, environment, SDK, and baselines) is open-sourced, and a challenge facing interdisciplinary researchers will be held based on the benchmark.

研究の動機と目的

  • 3D視覚入力からの操作における对象レベルの一般化性の評価を動機づけ、評価を可能にする。
  • 豊富なトポロジー・ジオメトリの変動を持つ多様な可動部品を提供し、クラス内一般化をテストする。
  • 異なる操作の課題を網羅する複数のタスクタイプを提供する(回転関節、直動、平面、制約のない動き)。
  • BC/オフラインRLのベースラインを促進する、成功した軌道の大規模データセットを用いた学習-from- demonstrations(LfD)をサポートする。
  • オープンで複数トラックのベンチマーク(視覚、RL、ロボティクス)を提供し、データ収集のスケーラビリティを確保して学際的な研究を促進する。)

提案手法

  • OpenCabinetDoor、OpenCabinetDrawer、PushChair、MoveBucketの4つの操作タスクを、 varied articulated objectsを用いて設計する。
  • ロボット搭載カメラからの自機視点パノラマ3D観測(点群、RGB-D)を用いて3D認識を可能にする。
  • RLベースのスケーラブルなパイプラインを介して約36,000件の成功デモ(約1.5Mの点群/RGB-Dフレーム)を収集し、共通報酬テンプレートとMPC支援検証を活用する。
  • 基準となる3Dディープラーニングポリシー(PointNet; PointNet + Transformer)とLfDアプローチ(模倣学習 BC; Offline RL BCQ, TD3+BC)を提供する。
  • 各タスク内で学習/テスト対象オブジェクトを分割し、複数のトラック(No Interactions、No External Annotations、No Restrictions)で对象レベルの一般化を評価する。
  • PostNet-Mobility資産を用い、凸分解・アーティファクト除去などの手動後処理と検証を行い、解ける環境を保証する。

実験結果

リサーチクエスチョン

  • RQ1多様なクラス内オブジェクト変異に跨って、3D視覚入力からのポリシーはオブジェクトレベルの一般化可能な操作スキルを学べるか。
  • RQ2PointNet、Transformerなどの3D深層学習アーキテクチャとLfD手法は、多様なオブジェクトで訓練した場合、オブジェクトレベルの一般化にどの程度適用できるか。
  • RQ3観察モダリティ(点群 vs RGB-D)がManiSkillの一般化性能に与える影響は何か。
  • RQ4デモがすべて成功である場合、オフラインRL手法は挙動模倣より優れているのか、どの条件下でそうなるのか。
  • RQ5回転関節、直動、平面、制約なしの異なるタスク運動は、ポリシー学習と一般化をどのように課題化するか。

主な発見

  • トポロジー/ジオメトリの大きなクラス内変動は、オブジェクトレベルの一般化性を評価するのに有効である。
  • デモがあっても全体の一般化性能は依然として難しく、訓練/テストの性能ギャップはタスク全体で顕著である。
  • BCを含むベースラインの中では、PointNet + Transformerがオブジェクトレベルの一般化に最も良いが、平均的なテスト成功率は依然として控えめ。
  • 提供デモに対してオフラインRL手法(BCQ、TD3+BC)は一貫してBCを上回るとは限らず、データとタスクの複雑さを浮き彫りにしている。
  • デモ数の増加は性能を向上させる;より多くの軌道がより高い成功率につながるが、未知のオブジェクトでは一般化は依然として非自明である。
  • 3次元入力戦略(セグメンテーションマスクを含む点群)とロボット状態の結合は、知覚とポリシー学習の重要な設計上の選択である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。