Skip to main content
QUICK REVIEW

[論文レビュー] Finite-Time Convergence Rates of Nonlinear Two-Time-Scale Stochastic Approximation under Markovian Noise

Thinh T. Doan|arXiv (Cornell University)|Apr 4, 2021
Markov Chains and Monte Carlo Methods参考文献 42被引用数 4
ひとこと要約

本稿は、マルコフノイズ下での非線形2時刻スケール確率的近似における有限時間収束レートを確立し、平均二乗誤差の期待値において ${\cal O}(1/k^{2/3})$ の収束速度を証明している。解析はリャプノフ関数と幾何的混合時間を用いて、従属データとバイアスを扱い、特異摂動理論を非漸近的設定に拡張し、マルコフ駆動サンプリングを伴うものに適用している。

ABSTRACT

We study the so-called two-time-scale stochastic approximation, a simulation-based approach for finding the roots of two coupled nonlinear operators. Our focus is to characterize its finite-time performance in a Markov setting, which often arises in stochastic control and reinforcement learning problems. In particular, we consider the scenario where the data in the method are generated by Markov processes, therefore, they are dependent. Such dependent data result to biased observations of the underlying operators. Under some fairly standard assumptions on the operators and the Markov processes, we provide a formula that characterizes the convergence rate of the mean square errors generated by the method to zero. Our result shows that the method achieves a convergence in expectation at a rate $\mathcal{O}(1/k^{2/3})$, where $k$ is the number of iterations. Our analysis is mainly motivated by the classic singular perturbation theory for studying the asymptotic convergence of two-time-scale systems, that is, we consider a Lyapunov function that carefully characterizes the coupling between the two iterates. In addition, we utilize the geometric mixing time of the underlying Markov process to handle the bias and dependence in the data. Our theoretical result complements for the existing literature, where the rate of nonlinear two-time-scale stochastic approximation under Markovian noise is unknown.

研究の動機と目的

  • データがマルコフ過程によって生成される場合の非線形2時刻スケール確率的近似の有限時間収束性能を特徴づけること。
  • 確率的近似アルゴリズムにおいて、マルコフ的サンプリングに起因する従属的・バイアスのある観測の課題に対処すること。
  • 標準的な作用素およびマルコフ過程に関する仮定の下で、反復の平均二乗誤差に対する非漸近的収束速度を導出すること。
  • 特異摂動理論を、マルコフノイズが存在する有限時間解析に拡張すること。
  • ステップサイズの選択と混合時間の収束速度への影響を定量化すること。

提案手法

  • 2時刻スケール系における速い反復と遅い反復の結合を分析するためにリャプノフ関数を用いる。
  • 基礎となるマルコフ過程の幾何的混合時間を用いて、従属サンプルに起因するバイアスを定量化・制御する。
  • 速い反復と遅い反復のためのステップサイズ $\alpha_k$ と $\beta_k$ に対して $\beta_k \ll \alpha_k$ の時間スケール分離を採用する。
  • 両方の反復が固定点から離れる合成偏差を追跡するため、修正された誤差過程 $\hat{z}_k = \hat{x}_k + \hat{y}_k$ を導入する。
  • 定数 $B$、$\mu_F$、$\mu_G$、および混合時間 $\tau(\alpha_k)$ を用いて誤差項の再帰的不等式とバウンドを適用し、収束バウンドを導出する。
  • 誤差項の成長を制御し、期待値の二乗誤差に対する一様バウンドを導出するために、$w_k$ を用いた指数的重み付けを採用する。

実験結果

リサーチクエスチョン

  • RQ1データがマルコフ過程によって生成される場合、非線形2時刻スケール確率的近似の有限時間収束レートは何か?
  • RQ2マルコフ過程の幾何的混合時間は、アルゴリズムのバイアスと収束にどのように影響するか?
  • RQ3特異摂動フレームワークは、マルコフ的サンプリング下で非漸近的収束速度を提供するために拡張可能か?
  • RQ4どのステップサイズの選択が、従属観測を処理しつつ最適な収束速度を保証するか?
  • RQ5速い反復と遅い反復の結合および非i.i.d.データに起因するバイアスが、平均二乗誤差にどのように共同で影響を与えるか?

主な発見

  • 本手法は、反復の平均二乗誤差の期待値において、有限時間収束レート ${\cal O}(1/k^{2/3})$ を達成している。
  • 収束速度は、作用素およびマルコフ過程に関する標準的仮定の下で導出されており、幾何的エルゴード性を含む。
  • マルコフ連鎖の幾何的混合時間は、従属サンプルに起因するバイアスを明示的に制御するために使用されている。
  • 解析により、誤差バウンドが $\alpha_k\beta_k$、$\alpha_k\alpha_{k;\tau(\alpha_k)}$、および $\beta_k^2$ の積に依存することが示され、係数は作用素のリプシッツ定数および固定点のノルムに依存する。
  • リャプノフ関数アプローチは、2つの時間スケール間の相互作用を的確に捉え、安定性と収束を保証している。
  • 本結果は、マルコフノイズ下での非線形2時刻スケールSAに対して、初めての有限時間収束レートを提供するという、文献における重要な空白を埋めている。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。