Skip to main content
QUICK REVIEW

[論文レビュー] The ODE Method for Asymptotic Statistics in Stochastic Approximation and Reinforcement Learning

Vivek S. Borkar, Shuhang Chen|arXiv (Cornell University)|Oct 27, 2021
Neural Networks and Applications被引用数 9
ひとこと要約

本稿は、Donsker-Varadhan (DV3) 条件の下で、新たなリャプノフ関数を用いて、基礎となるマルコフ過程の幾何的エルゴード性のもとで、確率的近似および強化学習アルゴリズムの漸近正規性を確立する。反復値およびその平均化バージョンの両方について、関数的およびスカラーの中心極限定理を証明し、正規化された共分散行列がPolyak-Ruppert平均化の最小漸近共分散に収束することを示す。

ABSTRACT

The paper concerns the $d$-dimensional stochastic approximation recursion, $$ θ_{n+1}= θ_n + α_{n + 1} f(θ_n, Φ_{n+1}) $$ where $ \{ Φ_n \}$ is a stochastic process on a general state space, satisfying a conditional Markov property that allows for parameter-dependent noise. The main results are established under additional conditions on the mean flow and a version of the Donsker-Varadhan Lyapunov drift condition known as (DV3): (i) An appropriate Lyapunov function is constructed that implies convergence of the estimates in $L_4$. (ii) A functional central limit theorem (CLT) is established, as well as the usual one-dimensional CLT for the normalized error. Moment bounds combined with the CLT imply convergence of the normalized covariance $ extsf{E}[ z_n z_n^T ]$ to the asymptotic covariance in the CLT, where $z_n =: (θ_n-θ^*)/\sqrt{α_n}$. (iii) The CLT holds for the normalized version $z^{ ext{PR}}_n =: \sqrt{n} [θ^{ ext{PR}}_n -θ^*]$, of the averaged parameters $θ^{ ext{PR}}_n =:n^{-1} \sum_{k=1}^nθ_k$, subject to standard assumptions on the step-size. Moreover, the covariance in the CLT coincides with the minimal covariance of Polyak and Ruppert. (iv) An example is given where $f$ and $\bar{f}$ are linear in $θ$, and $Φ$ is a geometrically ergodic Markov chain but does not satisfy (DV3). While the algorithm is convergent, the second moment of $θ_n$ is unbounded and in fact diverges. This arXiv version represents a major extension of the results in prior versions.The main results now allow for parameter-dependent noise, as is often the case in applications to reinforcement learning.

研究の動機と目的

  • ノイズ過程が幾何的エルゴード性を示すマルコフ連鎖である場合に、確率的近似アルゴリズムの漸近正規性を確立すること。
  • ODE法の枠組みのもとで、反復値およびその平均化バージョンの両方について、関数的およびスカラーの中心極限定理を証明すること。
  • 正規化された共分散行列がPolyak-Ruppert平均化の最小漸近共分散に収束することを示すこと。
  • 2次モーメントが有界のまま保たれる条件を特定すること、すなわちマルコフ連鎖がDV3条件を満たさない場合でも。

提案手法

  • Donsker-Varadhan (DV3) 条件の下で、連続過程 $(\theta_n, \Phi_n)$ のためのリャプノフ関数を構築し、$\theta_n$ の$L^4$収束を証明する。
  • 正規化誤差 $z_n = (\theta_n - \theta^*) / \sqrt{\alpha_n}$ に関数的中心極限定理 (CLT) の技法を適用し、拡散過程への弱収束を確立する。
  • 正規化誤差の漸近共分散行列 $\Sigma_\theta$ を導出し、その期待値における収束を証明する。
  • 平均化パラメータ $\tilde{\theta}^{\text{PR}}_n = n^{-1} \sum_{k=1}^n \theta_k$ を分析し、$\sqrt{n}(\theta^{\text{PR}}_n - \theta^*)$ に対するCLTを証明し、最小共分散 $\Sigma^{\text{PR}}_\theta$ を得る。
  • スケーリングされたODEとマルティングール差分分解を用いて、補間過程とODE解との間の誤差を制御する。
  • モーメントバウンドと大偏差を用いて、DV3条件が成立しない場合でも、発散する2次モーメントを確立する、すなわち収束が成立する場合でさえ。

実験結果

リサーチクエスチョン

  • RQ1ノイズ過程が幾何的エルゴード性を示すマルコフ連鎖である場合に、確率的近似再帰が $L^4$ で収束する条件は何か?
  • RQ2正規化誤差 $z_n = (\theta_n - \theta^*) / \sqrt{\alpha_n}$ は関数的中心極限定理を満たすか?
  • RQ3正規化された共分散 $\mathbb{E}[z_n z_n^\top]$ は漸近共分散 $\Sigma_\theta$ に収束するか?
  • RQ4平均化パラメータ $\tilde{\theta}^{\text{PR}}_n$ は最小漸近共分散 $\Sigma^{\text{PR}}_\theta$ を持つCLTを満たすか?
  • RQ52次モーメント $\mathbb{E}[\|\theta_n\|^2]$ が発散する場合でも、アルゴリズムが平均二乗収束する可能性はあるか?

主な発見

  • DV3条件の下で、連続過程 $(\theta_n, \Phi_n)$ のための構築されたリャプノフ関数により、$\theta_n$ の$L^4$収束が確立される。
  • 正規化誤差 $z_n = (\theta_n - \theta^*) / \sqrt{\alpha_n}$ に対して関数的中心極限定理が成立し、拡散過程への弱収束が成立する。
  • 正規化された共分散 $\mathbb{E}[z_n z_n^\top]$ は $n \to \infty$ のとき漸近共分散 $\Sigma_\theta$ に収束する。
  • 平均化パラメータ $\tilde{\theta}^{\text{PR}}_n$ に対してCLTが成立し、$\sqrt{n}(\theta^{\text{PR}}_n - \theta^*)$ は分布収束する。また、$\mathbb{E}[z^{\text{PR}}_n (z^{\text{PR}}_n)^\top]$ は $\Sigma^{\text{PR}}_\theta$、すなわちPolyak-Ruppert平均化の最小共分散に収束する。
  • $f$ と $\overline{f}$ が線形であり、マルコフ連鎖が幾何的エルゴード的であるが、DV3条件が成立しない例を構築した。その結果、$\mathbb{E}[\|\theta_n\|^2] \to \infty$ となるが、$\theta_n$ はほとんど確実に収束する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。