Skip to main content
QUICK REVIEW

[論文レビュー] Learning Agile Robotic Locomotion Skills by Imitating Animals

Xue Bin Peng, Erwin Coumans|arXiv (Cornell University)|Apr 2, 2020
Robotic Locomotion and Control被引用数 41
ひとこと要約

本論文は、実動物の動作を模倣することで、脚部ロボットが機敏な機動技術を習得する imitation-learning フレームワークを提示する。シミュレーションでドメインランダマイゼーションを用いて訓練し、 latent-space adaptation を介して現実のロボットへ転写する。

ABSTRACT

Reproducing the diverse and agile locomotion skills of animals has been a longstanding challenge in robotics. While manually-designed controllers have been able to emulate many complex behaviors, building such controllers involves a time-consuming and difficult development process, often requiring substantial expertise of the nuances of each skill. Reinforcement learning provides an appealing alternative for automating the manual effort involved in the development of controllers. However, designing learning objectives that elicit the desired behaviors from an agent can also require a great deal of skill-specific expertise. In this work, we present an imitation learning system that enables legged robots to learn agile locomotion skills by imitating real-world animals. We show that by leveraging reference motion data, a single learning-based approach is able to automatically synthesize controllers for a diverse repertoire behaviors for legged robots. By incorporating sample efficient domain adaptation techniques into the training process, our system is able to learn adaptive policies in simulation that can then be quickly adapted for real-world deployment. To demonstrate the effectiveness of our system, we train an 18-DoF quadruped robot to perform a variety of agile behaviors ranging from different locomotion gaits to dynamic hops and turns.

研究の動機と目的

  • 学習ベースの方法によってロボットに動物に似た敏捷性を達成するという課題を動機づけ、形式化する。
  • 手作業でスキル固有の報酬を設計することなく、実際の動物のモーションデータを活用してポリシー学習を導く。
  • ドメインランダマイゼーションと潜在空間適応によるサンプル効率の良いシムツーリアル転送を開発する。
  • 18-DoFの四脚歩行体で多様な機敏な挙動の学習を実証し、現実ロボットへ転送する。

提案手法

  • 逆運動学を用いて動物のモーションクリップをロボット形状へリターゲットする。
  • ゴール条件付き入力でリターゲット済みモーションを再現するよう、シミュレーションでモーション模倣ポリシーを訓練する。
  • PD制御トルク出力と姿勢/速度ベースの報酬を用いて参照軌道と一致させる。
  • 訓練中にさまざまなダイナミクスにポリシーを露出させるためにドメインランダマイゼーションを組み込む。
  • ランダム化されたダイナミクスを表す潜在 z にポリシーを条件づける潜在ダイナミクスエンコーダを導入し、ロバスト性と適応性のトレードオフを取る情報ボトルネックを設ける。
  • 実機で少数の試行で潜在エンコーディング z を適応させるために、アドバンテージ加重回帰に基づく適応手順(AWR)を適用する。

実験結果

リサーチクエスチョン

  • RQ1実際の動物モーションを用いて、脚式ロボットに対して頑健で多様な機 locomotion スキルを効果的に訓練できるか。
  • RQ2モーション模倣と潜在空間適応によって導かれた場合、ダイナミックな歩法全般でシムツーリアル転送は成立するか。
  • RQ3潜在ダイナミクスエンコーディングに情報ボトルネックを課すことが、実世界展開におけるロバスト性と適応性にどのように影響するか。
  • RQ4非適応または非ランダム化ベースラインと比べた場合、ドメインランダマイゼーションと潜在空間適応の組み合わせが実世界の性能に与える影響はどの程度か。

主な発見

  • ポリシーは18-DoF Laikago quadruped上で、歩行、跳躍、旋回など多様な機敏なスキルに適応する。
  • 適応型ポリシーは、現実のロボットへ転送した場合、ほとんどのスキルで非適応ベースラインを上回る。
  • 適応的手法は、Dog PaceやDog Spinなどのダイナミックなスキルを、頑丈だが非適応的なポリシーよりも信頼性高く実行できる。
  • 訓練は ~200 million simulation samples と ~50 real-world trials を各ポリシーで用い、挙動を適応させる。
  • 潜在空間適応で訓練されたポリシーは、非適応ポリシーよりも未知の動的環境の範囲に一般化する。
  • 後向き模倣モーションキャプチャデータ(例: dog gait)は、メーカーの歩行よりリアルワールドの速度を速く得られる(例: 1.08 m/s vs 0.84 m/s)。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。