Skip to main content
QUICK REVIEW

[論文レビュー] Model Identification via Physics Engines for Improved Policy Search.

Shaojun Zhu, Andrew Kimmel|arXiv (Cornell University)|Oct 24, 2017
Robot Manipulation and Learning被引用数 10
ひとこと要約

本論文では、市販の物理エンジンとブラックボックス型ベイズ最適化を組み合わせることで、ロボットシステムにおける質量や摩擦などの機械的パラメータを同定する手法を提案する。これにより、追加の現実世界データを必要とせず、実世界の軌道に近いシミュレーションモデルを構築し、正確なシミュレーションベースのポリシー探索を可能にする。この手法により、TRPO、PoWER、PILCOなどのアルゴリズムのポリシー学習性能が向上し、現実世界の軌道とよく一致する「十分に良い」モデルを構築する。

ABSTRACT

This paper presents a practical approach for identifying unknown mechanical parameters, such as mass and friction models of manipulated rigid objects or actuated robotic links, in a succinct manner that aims to improve the performance of policy search algorithms. Key features of this approach are the use of off-the-shelf physics engines and the adaptation of a black-box Bayesian optimization framework for this purpose. The physics engine is used to reproduce in simulation experiments that are performed on a real robot, and the mechanical parameters of the simulated system are automatically fine-tuned so that the simulated trajectories match with the real ones. The optimized model is then used for learning a policy in simulation, before safely deploying it on the real robot. Given the well-known limitations of physics engines in modeling real-world objects, it is generally not possible to find a mechanical model that reproduces in simulation the real trajectories exactly. Moreover, there are many scenarios where a near-optimal policy can be found without having a perfect knowledge of the system. Therefore, searching for a perfect model may not be worth the computational effort in practice. The proposed approach aims then to identify a model that is good enough to approximate the value of a locally optimal policy with a certain confidence, instead of spending all the computational resources on searching for the most accurate model. Empirical evaluations, performed in simulation and on a real robotic manipulation task, show that model identification via physics engines can significantly boost the performance of policy search algorithms that are popular in robotics, such as TRPO, PoWER and PILCO, with no additional real-world data.

研究の動機と目的

  • 不完全な物理エンジンのシミュレーションによる不正確なシステムモデルの影響を、ロボットのポリシー学習において解消すること。
  • ポリシー最適化に十分なモデルを対象とすることで、高精度なシステム同定にかかる計算コストを低減すること。
  • 校正済みのモデルを用いたシミュレーションでの学習により、現実のロボットへの安全かつ効率的なポリシーのデプロイを可能にすること。
  • TRPO、PoWER、PILCOのようなサンプル効率の良いポリシー探索アルゴリズムの性能を、モデル同定によって向上させること。
  • 完全なシステムモデルがなくても、シミュレーションが現実に十分に近い場合、近似的に最適なポリシーを学習可能であることを示すこと。

提案手法

  • 物理実験から得た現実世界のロボットの軌道を、モデルキャリブレーションの基準データとして活用する。
  • 物理エンジンを用いて、同じタスクをシミュレートし、異なる機械的パラメータのもとでの仮想軌道を生成する。
  • ブラックボックス型ベイズ最適化フレームワークを適用し、未知のパラメータ(例:質量、摩擦)を調整することで、シミュレートされた軌道と実際の軌道を一致させる。
  • モデルの完全な正確性を求めるのでなく、局所的に最適なポリシーを学習できるだけの十分な忠実度を求める。
  • 校正済みのシミュレーション環境でポリシーを学習した後、それを現実のロボットにデプロイする。
  • 校正済みのモデルを活用し、TRPO、PoWER、PILCOなどのポリシー探索アルゴリズムのサンプル効率と性能を向上させる。

実験結果

リサーチクエスチョン

  • RQ1市販の物理エンジンを用いて、最小限の現実世界データでロボットシステムの機械的パラメータを同定できるか?
  • RQ2完全なモデルではなく「十分に良い」モデルを同定することで、実際のポリシー学習性能が向上するか?
  • RQ3シミュレーションを用いたモデル同定は、TRPO や PILCO などのポリシー探索アルゴリズムの性能をどの程度向上させるか?
  • RQ4計算コストとポリシー性能の観点から、本手法は他の代替手法と比べてどのように差がつくか?
  • RQ5校正済みのシミュレーションモデルは、現実のロボットシステムへのポリシーの安全なデプロイを信頼できてか?

主な発見

  • 提案手法によるモデル同定は、シミュレーションおよび現実世界のロボット操作タスクにおいて、TRPO、PoWER、PILCO などのポリシー探索アルゴリズムの性能を顕著に向上させる。
  • 初期の軌道データ以外に追加の現実世界データを一切必要とせず、より高いポリシー学習性能を達成している。
  • 物理エンジンが本質的に完全な正確性を持たない場合でも、シミュレーションモデルを現実のダイナミクスに近づけるキャリブレーションに成功している。
  • 完全なモデルを同定するのと比べて、はるかに少ない計算コストで、ポリシー最適化に十分な「十分に良い」モデルを同定可能であることが実証された。
  • 校正済みのシミュレーション環境により、現実のロボットへのポリシーの安全かつ効果的なデプロイが可能となり、現実世界での実行時の失敗リスクが低減された。
  • 実験的結果により、本手法が複雑なロボット操作タスクにおけるポリシー探索アルゴリズムのサンプル効率と収束速度を向上させることを確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。