Skip to main content
QUICK REVIEW

[論文レビュー] DREAM Architecture: a Developmental Approach to Open-Ended Learning in Robotics

Stéphane Doncieux, Nicolas Bredèche|arXiv (Cornell University)|May 13, 2020
Reinforcement Learning in Robotics被引用数 10
ひとこと要約

DREAMアーキテクチャは、操作と学習に加えて第三の学習サイクル「再記述」を導入する発達的ロボティクスフレームワークを提案する。これにより、ロボットは開けた状態でタスクに依存しない学習を自律的に行うための状態表現を生成・適応可能となる。内的動機付け、トランスファーラーニング、睡眠に似たオフラインリプレイを統合することで、DREAMは連続的かつ現実世界の環境において安定した世界モデル学習と効率的なポリシー習得を可能にし、ナビゲーションタスクにおける耐障害性と表現の適応性の向上が実証された。

ABSTRACT

Robots are still limited to controlled conditions, that the robot designer knows with enough details to endow the robot with the appropriate models or behaviors. Learning algorithms add some flexibility with the ability to discover the appropriate behavior given either some demonstrations or a reward to guide its exploration with a reinforcement learning algorithm. Reinforcement learning algorithms rely on the definition of state and action spaces that define reachable behaviors. Their adaptation capability critically depends on the representations of these spaces: small and discrete spaces result in fast learning while large and continuous spaces are challenging and either require a long training period or prevent the robot from converging to an appropriate behavior. Beside the operational cycle of policy execution and the learning cycle, which works at a slower time scale to acquire new policies, we introduce the redescription cycle, a third cycle working at an even slower time scale to generate or adapt the required representations to the robot, its environment and the task. We introduce the challenges raised by this cycle and we present DREAM (Deferred Restructuring of Experience in Autonomous Machines), a developmental cognitive architecture to bootstrap this redescription process stage by stage, build new state representations with appropriate motivations, and transfer the acquired knowledge across domains or tasks or even across robots. We describe results obtained so far with this approach and end up with a discussion of the questions it raises in Neuroscience.

研究の動機と目的

  • 未知の環境で事前にプログラムされた状態や行動表現がなくとも、予期しないタスクを解けるロボットの実現に向けた課題に取り組む。
  • 従来の強化学習の限界を克服するため、状態表現を動的に生成・適応する「再記述サイクル」を導入する。
  • 共有可能で学習可能な表現を通じて、タスク間およびロボット間の知識転送を可能にする。
  • 人間の認知発達に類似した発達的プロセス、特に表現の再記述とオフライン学習をモデル化する。
  • ロボティクスと神経科学を橋渡しし、睡眠関連のリプレイと海馬機能に関する仮説を現実世界の学習において検証する。

提案手法

  • 時間スケールが徐々に遅くなる3サイクルフレームワーク(運用:ポリシー実行、学習:ポリシー最適化、再記述:表現生成・適応)を導入する。
  • 外部報酬が存在しない状況で、探索と表現発見を促進する内的動機付けメカニズムを採用する。
  • 過去に習得した知識をタスクや環境間で活用することで、サンプルの複雑さを低減するトランスファーラーニングを用いる。
  • 過去の経験をランダムな順序でオフラインリプレイすることで、連続的状態空間における学習の安定性を向上させ、睡眠中の海馬のリプレイを模倣する。
  • 反復的抽象化を通じて、低レベルのセンサモータデータを高レベルでタスクに適した表現に変換する表現の再記述を実施する。
  • 段階的な認知能力発達を可能にする、モジュール型でエンドツーエンドのアーキテクチャを統合する。

実験結果

リサーチクエスチョン

  • RQ1環境やタスクの事前知識がなくとも、ロボットが新しいタスクのために自律的かつ状態表現を生成・適応できる仕組みは何か?
  • RQ2時間的相関がオンライン学習を妨げる連続的かつ高次元の状態空間において、どのように安定した世界モデル学習が達成できるか?
  • RQ3内的動機付けとオフラインリプレイプロセスは、タスクやロボット間で再利用可能で転送可能な知識の習得をどのように支援するか?
  • RQ4ロボットの発達的プロセスが、人間の類似表現の再記述と認知発達をどの程度模倣できるか?
  • RQ5睡眠に似たオフラインプロセスは、現実世界のロボットシステムにおける強化学習の安定化とブートストラップにどのような役割を果たすか?

主な発見

  • 経験をランダムな順序でオフラインリプレイすることで、連続的状態空間ナビゲーションにおいて世界モデルの安定性と整合性が顕著に向上し、オンライン学習における時間的相関の問題を克服した。
  • 再記述サイクルにより、生のセンサモータデータが高レベルでタスクに適した表現に変換され、ポリシー学習の高速化と効率化が実現した。
  • 内的動機付けメカニズムが、情報量の多い状態へ向かう探索を導き、有用な表現や行動の発見を加速した。
  • タスク間のトランスファーラーニングにより、実世界でのサンプル数が削減され、学習済み表現の再利用性が実証された。
  • 睡眠に似たリプレイプロセスは、価値関数学習と世界モデル習得のブートストラップを支援し、哺乳類における提案された海馬機能を模倣した。
  • 本アーキテクチャは、現実世界のロボットタスクにおいて、純粋にエンドツーエンドまたはシミュレーションオンリーなアプローチに比べ、適応性とサンプル効率の面で優れた性能を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。