Skip to main content
QUICK REVIEW

[논문 리뷰] The ODE Method for Asymptotic Statistics in Stochastic Approximation and Reinforcement Learning

Vivek S. Borkar, Shuhang Chen|arXiv (Cornell University)|2021. 10. 27.
Neural Networks and Applications인용 수 9
한 줄 요약

이 논문은 Donsker-Varadhan (DV3) 조건 하에서, 기저가 되는 마르코프 과정이 기하적 에르고딕성을 만족할 경우, 확률적 근사 및 강화 학습 알고리즘의 점근적 정규성을 확립한다. 이는 새로운 리아푸노프 함수를 사용하여 이루어지며, 반복값과 그 평균화된 형태에 대해 기능적 및 스칼라 중심극한정리(central limit theorem)를 증명한다. 또한 정규화된 공분산 행렬이 Polyak-Ruppert 평균화의 최소 점근 공분산으로 수렴하는 것을 보여준다.

ABSTRACT

The paper concerns the $d$-dimensional stochastic approximation recursion, $$ θ_{n+1}= θ_n + α_{n + 1} f(θ_n, Φ_{n+1}) $$ where $ \{ Φ_n \}$ is a stochastic process on a general state space, satisfying a conditional Markov property that allows for parameter-dependent noise. The main results are established under additional conditions on the mean flow and a version of the Donsker-Varadhan Lyapunov drift condition known as (DV3): (i) An appropriate Lyapunov function is constructed that implies convergence of the estimates in $L_4$. (ii) A functional central limit theorem (CLT) is established, as well as the usual one-dimensional CLT for the normalized error. Moment bounds combined with the CLT imply convergence of the normalized covariance $ extsf{E}[ z_n z_n^T ]$ to the asymptotic covariance in the CLT, where $z_n =: (θ_n-θ^*)/\sqrt{α_n}$. (iii) The CLT holds for the normalized version $z^{ ext{PR}}_n =: \sqrt{n} [θ^{ ext{PR}}_n -θ^*]$, of the averaged parameters $θ^{ ext{PR}}_n =:n^{-1} \sum_{k=1}^nθ_k$, subject to standard assumptions on the step-size. Moreover, the covariance in the CLT coincides with the minimal covariance of Polyak and Ruppert. (iv) An example is given where $f$ and $\bar{f}$ are linear in $θ$, and $Φ$ is a geometrically ergodic Markov chain but does not satisfy (DV3). While the algorithm is convergent, the second moment of $θ_n$ is unbounded and in fact diverges. This arXiv version represents a major extension of the results in prior versions.The main results now allow for parameter-dependent noise, as is often the case in applications to reinforcement learning.

연구 동기 및 목표

  • 소음 과정이 기하적 에르고딕성을 만족하는 마르코프 체인이 되는 경우, 확률적 근사 알고리즘의 점근적 정규성을 확립하는 것.
  • ODE 방법 프레임워크 하에서 반복값과 그 평균화된 형태에 대해 기능적 및 스칼라 중심극한정리를 증명하는 것.
  • 정규화된 공분산 행렬이 Polyak-Ruppert 평균화의 최소 점근 공분산으로 수렴하는 것을 보이는 것.
  • 두 번째 모멘트가 유한하게 유지되는 조건을 특정하는 것 — 이는 마르코프 체인이 DV3 조건을 만족하지 않을 경우에도 해당된다.

제안 방법

  • Donsker-Varadhan (DV3) 조건 하에서 연합 과정 frac{\theta_n, \Phi_n}에 대해 리아푸노프 함수를 구성하여 \theta_n의 L^4 수렴성을 증명한다.
  • 정규화된 오차 z_n = (\theta_n - \theta^*) / \sqrt{\alpha_n}에 기능적 중심극한정리 기법을 적용하여, 강한 수렴이 확산으로 수렴함을 증명한다.
  • 정규화된 오차의 점근 공분산 행렬 \Sigma_\theta를 유도하고, 기대값 수렴을 증명한다.
  • 평균화된 매개변수 \tilde{\theta}^{\text{PR}}_n = n^{-1} \sum_{k=1}^n \theta_k를 분석하여, \sqrt{n}(\theta^{\text{PR}}_n - \theta^*)에 대해 중심극한정리를 증명하고, 최소 공분산 \Sigma^{\text{PR}}_\theta를 얻는다.
  • 스케일된 ODE와 민트링글 차분 분해를 사용하여, 보간 과정과 ODE 해 사이의 오차를 통제한다.
  • 모멘트 경계와 대규모 편차 이론을 활용하여, DV3 조건이 실패할 경우에도 정규화된 두 번째 모멘트가 유계가 아님을 증명한다 — 이는 수렴 조건 하에서도 마찬가지이다.

실험 결과

연구 질문

  • RQ1소음 과정이 기하적 에르고딕성을 만족하는 마르코프 체인이 되는 경우, 확률적 근사 재귀식이 L^4 수렴성을 보일 조건은 무엇인가?
  • RQ2정규화된 오차 z_n = (\theta_n - \theta^*) / \sqrt{\alpha_n}이 기능적 중심극한정리를 만족하는가?
  • RQ3정규화된 공분산 \mathbb{E}[z_n z_n^\top]은 점근 공분산 \Sigma_\theta로 수렴하는가?
  • RQ4평균화된 매개변수 \tilde{\theta}^{\text{PR}}_n은 최소 점근 공분산 \Sigma^{\text{PR}}_\theta를 갖는 중심극한정리를 만족하는가?
  • RQ5두 번째 모멘트 \mathbb{E}[\|\theta_n\|^2]이 발산하더라도, 알고리즘이 평균 제곱 수렴을 이룰 수 있는가?

주요 결과

  • Donsker-Varadhan (DV3) 조건 하에서 연합 과정 (\theta_n, \Phi_n)에 대해 구성된 리아푸노프 함수를 통해 \theta_n의 L^4 수렴성이 확립된다.
  • 정규화된 오차 z_n = (\theta_n - \theta^*) / \sqrt{\alpha_n}에 대해 기능적 중심극한정리가 성립하며, 약한 수렴이 확산으로 수렴한다.
  • 정규화된 공분산 \mathbb{E}[z_n z_n^\top]은 n \to \infty 일 때 점근 공분산 \Sigma_\theta로 수렴한다.
  • 평균화된 매개변수 \tilde{\theta}^{\text{PR}}_n에 대해 중심극한정리가 성립하며, \sqrt{n}(\theta^{\text{PR}}_n - \theta^*)는 분포 수렴을 이루고, \mathbb{E}[z^{\text{PR}}_n (z^{\text{PR}}_n)^\top]은 Polyak-Ruppert 평균화의 최소 공분산 \Sigma^{\text{PR}}_\theta로 수렴한다.
  • f와 \overline{f}가 선형인 예시를 구성하였으며, 마르코프 체인이 기하적 에르고딕성을 만족하지만 DV3 조건은 실패한다. 이 경우 \mathbb{E}[\|\theta_n\|^2] \to \infty 이지만 \theta_n는 거의 확실히 수렴한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.