Skip to main content
QUICK REVIEW

[論文レビュー] Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity

Laixi Shi, Yuejie Chi|arXiv (Cornell University)|Aug 11, 2022
Reinforcement Learning in Robotics被引用数 8
ひとこと要約

本稿では、分布的に頑健な価値反復とデータ駆動型の悲観的ペナルティを組み合わせることで、近似的に最適なサンプル複雑度を達成する、モデルベースのオフライン強化学習アルゴリズムを提案する。軽微な分布シフト仮定の下で有限サンプル性能保証を確立し、(有効な)ホライズン長の多項式要因を除いて境界のタイトさを証明する。

ABSTRACT

This paper concerns the central issues of model robustness and sample efficiency in offline reinforcement learning (RL), which aims to learn to perform decision making from history data without active exploration. Due to uncertainties and variabilities of the environment, it is critical to learn a robust policy -- with as few samples as possible -- that performs well even when the deployed environment deviates from the nominal one used to collect the history dataset. We consider a distributionally robust formulation of offline RL, focusing on tabular robust Markov decision processes with an uncertainty set specified by the Kullback-Leibler divergence in both finite-horizon and infinite-horizon settings. To combat with sample scarcity, a model-based algorithm that combines distributionally robust value iteration with the principle of pessimism in the face of uncertainty is proposed, by penalizing the robust value estimates with a carefully designed data-driven penalty term. Under a mild and tailored assumption of the history dataset that measures distribution shift without requiring full coverage of the state-action space, we establish the finite-sample complexity of the proposed algorithms. We further develop an information-theoretic lower bound, which suggests that learning RMDPs is at least as hard as the standard MDPs when the uncertainty level is sufficient small, and corroborates the tightness of our upper bound up to polynomial factors of the (effective) horizon length for a range of uncertainty levels. To the best our knowledge, this provides the first provably near-optimal robust offline RL algorithm that learns under model uncertainty and partial coverage.

研究の動機と目的

  • モデルの不確実性と限られたデータの下で、頑健な方策を学習する課題に対処すること。
  • 完全な状態-行動空間カバレッジを必要とせずに、分布シフトと部分的カバレッジを克服すること。
  • 有限ホライズンおよび無限ホライズンの両設定において、近似的に最適なサンプル複雑度を達成すること。
  • KLダイバージェンスに基づく不確実性集合を用いた分布的に頑健なMDPにおける性能保証を提供すること。
  • 情報理論的に最適な多項式要因を除いてタイトな有限サンプル境界を確立すること。

提案手法

  • 経験的遷移モデルの周囲にKLダイバージェンスに基づく不確実性集合を用いて、分布的に頑健なMDP(RMDP)を定式化する。
  • 不確実性に対処するため、データ駆動型ペナルティ項で価値推定をペナルティ化した、頑健価値反復の悲観的変種であるDRVI-LCBを提案する。
  • オフラインデータから経験的ノーマルMDPを構築し、分布シフトに対して頑健性を確保する。
  • 状態-行動訪問頻度の逆数に比例するスケーリングを行う、新しいデータ駆動型ペナルティ項を導入し、悲観的推定を強化する。
  • 濃度不等式とKLダイバージェンスの性質を用いて、方策の頑健価値関数における有限サンプル境界を導出する。
  • 情報理論的下界を証明することで、上界のタイトさが(有効な)ホライズン長の多項式要因を除いて、情報理論的に最適であることを示す。

実験結果

リサーチクエスチョン

  • RQ1分布シフトに強く、かつサンプル効率の良いモデルベースのオフライン強化学習アルゴリズムを設計できるか?
  • RQ2データ不足に対処するため、悲観的アプローチを分布的に頑健な価値反復に効果的に統合する方法は何か?
  • RQ3部分的カバレッジを持つオフライン強化学習において、頑健な方策を学習するための根本的なサンプル複雑度の限界は何か?
  • RQ4提案されたアルゴリズムのサンプル複雑度は近似的に最適であり、情報理論的下界と比較してどう異なるか?
  • RQ5どのようなデータ仮定のもとで、アルゴリズムは有限サンプル性能保証を達成するか?

主な発見

  • 提案されたDRVI-LCBアルゴリズムは、完全なカバレッジを必要とせず、履歴データセットの分布シフトを測る軽微な仮定のもとで、有限サンプル性能保証を達成する。
  • アルゴリズムのサンプル複雑度は近似的に最適であり、上界が(有効な)ホライズン長の多項式要因を除いて情報理論的下界と一致する。
  • 情報理論的下界は、不確実性レベルが小さいとき、RMDPを学習することは標準MDPを学習するのと同程度以上に難しいことを示しており、上界のタイトさを確認する。
  • データ駆動型ペナルティ項は効果的に悲観的推定を強化し、不確実性集合の事前知識がなくても、低データ環境での頑健性を向上させる。
  • 理論的分析により、行動方策が状態-行動空間を完全にカバーしていない場合でも、アルゴリズムが有効に機能することが確認された。
  • 結果は有限ホライズンおよび無限ホライズンの両設定に適用可能であり、広範な適用可能性を示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。