Skip to main content
QUICK REVIEW

[論文レビュー] DMAP: a Distributed Morphological Attention Policy for Learning to Locomote with a Changing Body

Alberto Silvio Chiappa, Alessandro Marin Vargas|arXiv (Cornell University)|Sep 28, 2022
Action Observation and Synchronization被引用数 8
ひとこと要約

DMAPは、体の形態パラメータへの明示的アクセスがなくても、多様な身体形態を持つエージェントが歩行を学習できる生物学的にインspiredな注目メカニズムを備えた深層強化学習アーキテクチャを提案する。分散関節制御、独立した本体認識処理、動的注目ゲーティングを組み合わせることで、4つの連続的制御環境において、形態パラメータを完全に把握したオラクルエージェントと同等またはそれ以上の性能を達成する。

ABSTRACT

Biological and artificial agents need to deal with constant changes in the real world. We study this problem in four classical continuous control environments, augmented with morphological perturbations. Learning to locomote when the length and the thickness of different body parts vary is challenging, as the control policy is required to adapt to the morphology to successfully balance and advance the agent. We show that a control policy based on the proprioceptive state performs poorly with highly variable body configurations, while an (oracle) agent with access to a learned encoding of the perturbation performs significantly better. We introduce DMAP, a biologically-inspired, attention-based policy network architecture. DMAP combines independent proprioceptive processing, a distributed policy with individual controllers for each joint, and an attention mechanism, to dynamically gate sensory information from different body parts to different controllers. Despite not having access to the (hidden) morphology information, DMAP can be trained end-to-end in all the considered environments, overall matching or surpassing the performance of an oracle agent. Thus DMAP, implementing principles from biological motor control, provides a strong inductive bias for learning challenging sensorimotor tasks. Overall, our work corroborates the power of these principles in challenging locomotion tasks.

研究の動機と目的

  • 大きな、予期しない身体形態の変化があっても、ロバストに歩行を学習できる強化学習ポリシーの開発。
  • 生物学的にインspiredな原則(分散制御、独立した感覚処理、動的ゲーティング)が、センサモータリ適応に強いインダクティブバイアスを提供するかどうかの調査。
  • 形態パラメータが隠されているにもかかわらず、エンドツーエンドで学習可能なポリシーが、オラクルエージェントの性能に達するか、それを上回るかの評価。
  • 注目メカニズムのダイナミクスと、形態の変化にわたる一般化におけるその役割の分析。

提案手法

  • DMAPは、各体部に対して独立して、重み共有の時系列畳み込みネットワーク(TCN)を用いて本体状態履歴を処理する。
  • 感覚入力から得られる値符号化ベクトルに、学習された注目重みを適用することで、関節ごとの形態符号化を計算する。
  • 各関節は、現在の本体認識とその関節固有の形態符号化を用いる独立した全結合ネットワークによって制御される。
  • 注目メカニズムは、タスクに必要な形態に応じて、異なる体部からの感覚情報を異なる制御器に動的にゲーティングする。
  • 全アーキテクチャは、形態パラメータの教師なしで、スパarsityな密集報酬のみに依存してエンドツーエンドで学習される。
  • 分布内(IID)および分布外(OOD)の形態的摂動の両方を用いて評価し、ゼロショット一般化をテストする。

実験結果

リサーチクエスチョン

  • RQ1各エピソードの開始時に、身体形態がランダムに摂動された場合でも、深層強化学習ポリシーが効果的に歩行を学習できるか?
  • RQ2生物学的運動制御の原則(分散制御、独立処理、動的ゲーティング)を組み込むことで、未観測の形態に一般化する能力が向上するか?
  • RQ3エンドツーエンドで学習されたエージェントが、隠れた形態パラメータにアクセスできるオラクルエージェントの性能に達するか、それを上回れるか?
  • RQ4DMAPの注目メカニズムは形態の変化にどのように適応するか?訓練中にどのようなダイナミクスが出現するか?

主な発見

  • DMAPは、形態パラメータにアクセスしていないにもかかわらず、4つの環境(Ant, Half Cheetah, Walker, Hopper)すべてでオラクルエージェントと同等またはそれ以上の性能を達成する。
  • 単純なベースライン(本体認識のみに依存)は、多様な形態において効果的な歩行戦略を学習できない。
  • DMAPは、四肢切断(100%の摂動)を含む分布外の形態的摂動に対しても、ロバストに一般化する。報酬は単調に減少するが、四肢の種類に関係なく安定を保つ。
  • 注目メカニズムは、埋め込み空間で回転ダイナミクスを示し、訓練の進行に伴い徐々に出現し、分離され、正則化された軌道をもたらす。
  • 観測および行動空間の軌道と比較して、注目ダイナミクスははるかに分離されており、一般化性およびロバスト性の向上を示唆する。
  • DMAPとRMAの間で、新しい形態への適応速度に差はなく、二段階の訓練手順を必要とせず、迅速な適応を達成している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。