Skip to main content
QUICK REVIEW

[論文レビュー] On Reward-Free RL with Kernel and Neural Function Approximations: Single-Agent MDP and Markov Game

Shuang Qiu, Jieping Ye|arXiv (Cornell University)|Oct 19, 2021
Reinforcement Learning in Robotics被引用数 4
ひとこと要約

本稿では、単一エージェントMDPおよびゼロサムマルコフゲームの両方において、カーネルおよびニューラル関数近似を用いた、初めての証明可能に効率的な報酬フリー強化学習アルゴリズムを提案する。関数近似器から得られる探索ボーナスを用いた楽観的価値反復を採用することで、任意の外生的報酬に対して、$ʹ\mathcal{O}}(1/\varepsilon^{2})$ のサンプル複雑度で $ʹ$-サブオプティマルな方策、または $ʹ$-近似ナッシュ均衡を達成する。

ABSTRACT

To achieve sample efficiency in reinforcement learning (RL), it necessitates efficiently exploring the underlying environment. Under the offline setting, addressing the exploration challenge lies in collecting an offline dataset with sufficient coverage. Motivated by such a challenge, we study the reward-free RL problem, where an agent aims to thoroughly explore the environment without any pre-specified reward function. Then, given any extrinsic reward, the agent computes the policy via a planning algorithm with offline data collected in the exploration phase. Moreover, we tackle this problem under the context of function approximation, leveraging powerful function approximators. Specifically, we propose to explore via an optimistic variant of the value-iteration algorithm incorporating kernel and neural function approximations, where we adopt the associated exploration bonus as the exploration reward. Moreover, we design exploration and planning algorithms for both single-agent MDPs and zero-sum Markov games and prove that our methods can achieve $\widetilde{\mathcal{O}}(1 /\varepsilon^2)$ sample complexity for generating a $\varepsilon$-suboptimal policy or $\varepsilon$-approximate Nash equilibrium when given an arbitrary extrinsic reward. To the best of our knowledge, we establish the first provably efficient reward-free RL algorithm with kernel and neural function approximators.

研究の動機と目的

  • 事前指定された報酬が存在しないオフライン強化学習における、サンプル効率の良い探索の課題に取り組む。
  • カーネルやニューラルネットワークのような非線形関数近似器をサポートする、報酬フリー強化学習の統一的枠組みを構築する。
  • 表形式および線形設定から、単一エージェントMDPおよびマルコフゲームにおける非線形関数近似への、証明可能に効率的なアルゴリズムの拡張を行う。
  • 任意の外生的報酬に対して、最適なサンプル複雑度を達成するように設計された、探索と計画のアルゴリズムを設計する。

提案手法

  • カーネルおよびニューラル関数近似器を用いた最小二乗価値反復の楽観的変種を提案し、内在的報酬として探索ボーナスを生成する。
  • 関数近似における関連する(スケーリングされた)ボーナスを、オフラインデータ収集段階での探索報酬として使用する。
  • 任意の与えられた外生的報酬と、オフラインデータのみを用いて方策またはナッシュ均衡を計算する計画段階を設計する。
  • マルコフゲームの場合、Q関数から形成される行列ゲームを解くことで計画段階を実装するが、計算上は効率的である。
  • カーネルおよびニューラル関数空間における集中不等式と正則化を活用し、推定誤差を抑え、一般化を保証する。
  • 正則化と信頼区間を用いた不確実性の定量化を統合し、非線形関数近似設定における探索を誘導する。

実験結果

リサーチクエスチョン

  • RQ1カーネルおよびニューラル関数近似を用いた、単一エージェントMDPにおける証明可能に効率的な報酬フリー強化学習アルゴリズムを設計できるか?
  • RQ2提案手法は、非線形関数近似設定において、$ʹ\mathcal{O}}(1/\varepsilon^{2})$ のサンプル複雑度を達成できるか?
  • RQ3このフレームワークは、類似したサンプル効率の保証を持つ、マルチエージェントゼロサムマルコフゲームへ拡張可能か?
  • RQ4外生的報酬が存在しない状況で、探索ボーナスを非線形関数近似器と整合的に効果的に設計する方法は何か?

主な発見

  • 提案アルゴリズムは、カーネルおよびニューラル関数近似を用いた単一エージェントMDPにおいて、$ʹ\mathcal{O}}(1/\varepsilon^{2})$ のサンプル複雑度で $ʹ$-サブオプティマルな方策を生成することを達成する。
  • ゼロサムマルコフゲームでは、$ʹ\mathcal{O}}(1/\varepsilon^{2})$ のサンプル複雑度で $ʹ$-近似ナッシュ均衡を計算する。
  • マルコフゲームにおける計画段階は、Q関数から形成される行列ゲームを解くものであり、計算的に効率的で、独立に価値のある結果である。
  • 理論的解析により、アルゴリズムの誤差バウンドが、高確率で $ʹ\mathcal{O}}\left(\beta\sqrt{H^{4}[\Gamma(K,\lambda;\ker_{m})+\log(KH)]}/\sqrt{K}+H^{2}\beta\iota\right)$ のスケーリングに従うことが示された。
  • 本手法は、カーネルおよびニューラルネットワークを含む非線形関数近似器を用いた報酬フリー強化学習において、初めての証明可能なサンプル効率を提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。