Skip to main content
QUICK REVIEW

[論文レビュー] Model-based Reinforcement Learning: A Survey

Thomas M. Moerland, Joost Broekens|arXiv (Cornell University)|Jun 30, 2020
Complex Systems and Decision Making被引用数 82
ひとこと要約

この論文はモデルベース強化学習を概説し、ダイナミクスモデルの学習と計画-学習の統合を詳述し、含意的なモデルベースRLと潜在的な利点を含む。

ABSTRACT

Sequential decision making, commonly formalized as Markov Decision Process (MDP) optimization, is a important challenge in artificial intelligence. Two key approaches to this problem are reinforcement learning (RL) and planning. This paper presents a survey of the integration of both fields, better known as model-based reinforcement learning. Model-based RL has two main steps. First, we systematically cover approaches to dynamics model learning, including challenges like dealing with stochasticity, uncertainty, partial observability, and temporal abstraction. Second, we present a systematic categorization of planning-learning integration, including aspects like: where to start planning, what budgets to allocate to planning and real data collection, how to plan, and how to integrate planning in the learning and acting loop. After these two sections, we also discuss implicit model-based RL as an end-to-end alternative for model learning and planning, and we cover the potential benefits of model-based RL. Along the way, the survey also draws connections to several related RL fields, like hierarchical RL and transfer learning. Altogether, the survey presents a broad conceptual overview of the combination of planning and learning for MDP optimization.

研究の動機と目的

  • モデルベースRLパラダイムにおける計画と学習の統合を説明する。
  • ダイナミクスモデルがどのように学習されるかを分類し、関与する課題(確率性、不確実性、部分観測性、非定常性)を整理する。
  • 計画と学習の統合方法と、学習ループ内でいつ計画を開始すべきかを系統的に論じる。
  • エンドツーエンドの代替として暗黙的モデルベースRLを紹介し、その潜在的利点を論じる。

提案手法

  • ダイナミクスへのアクセスとグローバル解の有無に基づいて、モデルベースRLを定義し、それを計画とモデルフリーRLと区別する。
  • 学習済みモデルを有するモデルベースRL、既知のモデルを有するモデルベースRL、学習済みモデル上での計画という三形態の計画-学習統合を提示する。
  • 正順前方モデル、後方モデル、逆モデルを含むダイナミクスモデル学習をレビューし、推定手法(parametric vs non-parametric, exact vs approximate)を論じる。
  • モデル有効域(global vs local)とこれが計画と学習に与える影響を論じる。
  • モデリングにおける課題:stochasticity、uncertainty、partial observability、non-stationarity、multi-step prediction、state and temporal abstractionを扱う。
  • 計画と学習の統合アプローチを概説する、計画開始状態、計画予算、計画手法、エンドツーエンド学習の派生を含む。

実験結果

リサーチクエスチョン

  • RQ1MDP最適化において、計画と学習をどのように統合して効果的なモデルベースRLアルゴリズムを形成できるか。
  • RQ2ダイナミクスモデル学習の主要な課題は何であり、それらはどのように対処できるか(stochasticity、uncertainty、partial observability、non-stationarity)?
  • RQ3global vs local model validityのトレードオフと、それらが計画とデータ効率に与える影響は何か?
  • RQ4学習ループ内で計画をどのように整理すべきか(いつ計画するのか、どれだけデータを集めるのか、どう計画するのか)?

主な発見

  • モデルベースRLはダイナミクスモデルとグローバル解を組み合わせてサンプル効率を改善する。
  • パラメトリック近似とグローバルカバレッジを持つforwardモデルは実務で一般的である。
  • 部分観測性と非定常性には、信念状態、再帰、部分モデルなどの特殊な技術が必要である。
  • 多ステップ予測の課題は、長期的な精度を向上させるために、多ステップ損失の使用や専用のn-stepモデルの導入を促す。
  • 計画は学習済みモデル上で行われ、方策/価値関数の学習、あるいはその両方が含まれる場合がある。学習されたモデル上での計画は、グローバル解が学習されていない場合、モデルベースRLとはみなされない。
  • 暗黙的モデルベースRLは、計画とモデリングの利点を活用しつつ、エンドツーエンドの代替手段を提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。