Skip to main content
QUICK REVIEW

[論文レビュー] Weak Convergence Properties of Constrained Emphatic Temporal-difference Learning with Constant and Slowly Diminishing Stepsize

Huizhen Yu|arXiv (Cornell University)|Nov 23, 2015
Reinforcement Learning in Robotics参考文献 36被引用数 12
ひとこと要約

本稿は、オフポリシー強化学習における定数ステップサイズおよび徐々に減少するステップサイズの両方を用いた制約付き強調時系列学習(ETD(λ))の弱収束特性を確立する。一般のオフポリシー条件下で、最適価値関数の近傍への漸近的収束を示し、弱Fellerマルコフ連鎖理論と確率的近似を用いて、固定定数ステップサイズにおける偏差の上限を導出する。

ABSTRACT

We consider the emphatic temporal-difference (TD) algorithm, ETD($λ$), for learning the value functions of stationary policies in a discounted, finite state and action Markov decision process. The ETD($λ$) algorithm was recently proposed by Sutton, Mahmood, and White to solve a long-standing divergence problem of the standard TD algorithm when it is applied to off-policy training, where data from an exploratory policy are used to evaluate other policies of interest. The almost sure convergence of ETD($λ$) has been proved in our recent work under general off-policy training conditions, but for a narrow range of diminishing stepsize. In this paper we present convergence results for constrained versions of ETD($λ$) with constant stepsize and with diminishing stepsize from a broad range. Our results characterize the asymptotic behavior of the trajectory of iterates produced by those algorithms, and are derived by combining key properties of ETD($λ$) with powerful convergence theorems from the weak convergence methods in stochastic approximation theory. For the case of constant stepsize, in addition to analyzing the behavior of the algorithms in the limit as the stepsize parameter approaches zero, we also analyze their behavior for a fixed stepsize and bound the deviations of their averaged iterates from the desired solution. These results are obtained by exploiting the weak Feller property of the Markov chains associated with the algorithms, and by using ergodic theorems for weak Feller Markov chains, in conjunction with the convergence results we get from the weak convergence methods. Besides ETD($λ$), our analysis also applies to the off-policy TD($λ$) algorithm, when the divergence issue is avoided by setting $λ$ sufficiently large.

研究の動機と目的

  • 標準TD(λ)が行動方策とターゲット方策が異なるオフポリシー学習で長年の発散問題を抱えることへの対処。
  • 定数ステップサイズおよび徐々に減少するステップサイズスケジュールの両方において、制約付きETD(λ)の収束保証の確立。
  • 確率的近似理論からの弱収束手法を用いて、ETD(λ)反復の漸近的挙動の特定。
  • λが十分に大きい場合に、これらの結果をオフポリシーTD(λ)へ拡張し、制約付きバージョンの新たな収束特性を提供。
  • 有限サンプル挙動と定数ステップサイズ設定における偏差の上限の分析により、実用的安定性と解釈可能性の向上。

提案手法

  • 確率的近似理論からの弱収束手法を適用し、制約付きETD(λ)反復の極限的挙動を分析。
  • 学習ダイナミクスによって誘導されるマルコフ連鎖の弱Feller性を用いて、エルゴード性と安定性を確立。
  • 平均ODE(常微分方程式)を導出し、平均化された反復の極限的挙動を記述。
  • オフポリシー設定における学習の安定化と分散低減を目的に、バイアス項を含むETD(λ)の2つの変種を導入。
  • 弱Feller連鎖のエルゴード定理を用いて、サンプルパス挙動と射影ベルマン方程式の解とを結びつける。
  • λが十分に大きい場合に、収束特性の同等性を示すことにより、結果をオフポリシーTD(λ)へ拡張。

実験結果

リサーチクエスチョン

  • RQ1定数ステップサイズを用いた制約付きETD(λ)が、最適価値関数の近傍に弱収束する条件は何か?
  • RQ2徐々に減少するステップサイズを用いたETD(λ)の漸近的挙動は、定数ステップサイズを用いた場合とどのように比較されるか?
  • RQ3固定定数ステップサイズ下で、平均化された反復の最適解からの偏差の上限は何か?
  • RQ4λが十分に大きい場合に、ETD(λ)の収束結果はオフポリシーTD(λ)へ拡張可能か?
  • RQ5学習ダイナミクスによって誘導されるマルコフ連鎖の弱Feller性は、制約付きETD(λ)の収束解析にどのように寄与するか?

主な発見

  • 定数ステップサイズを用いた制約付きETD(λ)は、最適解の近傍に弱収束し、その偏差はステップサイズに比例する項で有界である。
  • 徐々に減少するステップサイズの下では、一般のオフポリシー条件下で、アルゴリズムは弱収束して最適解に到達する。
  • 制約付きETD(λ)の平均化されたプロセスの平均ODEは、一意の全域的吸引平衡点を持つため、望ましい価値関数への収束が保証される。
  • λが十分に大きい(例:λ=1)場合、同じ収束結果がオフポリシーTD(λ)にも適用可能であり、ETD(λ)を超えた解析が可能になる。
  • 学習ダイナミクスによって誘導されるマルコフ連鎖の弱Feller性により、エルゴード性が保証され、収束解析にエルゴード定理の適用が可能になる。
  • 本分析により、特にオフポリシー設定における高分散重要度サンプリングの取り扱いにおいて、制約付きETD(λ)およびオフポリシーTD(λ)に対する新たな理論的保証が得られる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。