Skip to main content
QUICK REVIEW

[论文解读] Finite-Time Convergence Rates of Nonlinear Two-Time-Scale Stochastic Approximation under Markovian Noise

Thinh T. Doan|arXiv (Cornell University)|Apr 4, 2021
Markov Chains and Monte Carlo Methods参考文献 42被引用 4
一句话总结

本文在马尔可夫噪声下建立了非线性双时标随机逼近的有限时间收敛速率,证明了均方误差的期望收敛速率为 ${\cal O}(1/k^{2/3})$。分析利用李雅普诺夫函数与几何混合时间来处理依赖数据与偏差,将奇异摄动理论扩展至具有马尔可夫驱动采样的非渐近设置。

ABSTRACT

We study the so-called two-time-scale stochastic approximation, a simulation-based approach for finding the roots of two coupled nonlinear operators. Our focus is to characterize its finite-time performance in a Markov setting, which often arises in stochastic control and reinforcement learning problems. In particular, we consider the scenario where the data in the method are generated by Markov processes, therefore, they are dependent. Such dependent data result to biased observations of the underlying operators. Under some fairly standard assumptions on the operators and the Markov processes, we provide a formula that characterizes the convergence rate of the mean square errors generated by the method to zero. Our result shows that the method achieves a convergence in expectation at a rate $\mathcal{O}(1/k^{2/3})$, where $k$ is the number of iterations. Our analysis is mainly motivated by the classic singular perturbation theory for studying the asymptotic convergence of two-time-scale systems, that is, we consider a Lyapunov function that carefully characterizes the coupling between the two iterates. In addition, we utilize the geometric mixing time of the underlying Markov process to handle the bias and dependence in the data. Our theoretical result complements for the existing literature, where the rate of nonlinear two-time-scale stochastic approximation under Markovian noise is unknown.

研究动机与目标

  • 研究当数据由马尔可夫过程生成时,非线性双时标随机逼近的有限时间收敛性能。
  • 解决由于马尔可夫采样导致的依赖性与偏差观测在随机逼近算法中的挑战。
  • 在算子与马尔可夫过程的标准假设下,推导出迭代均方误差的非渐近收敛速率。
  • 将奇异摄动理论扩展至马尔可夫噪声存在下的有限时间分析。
  • 量化步长选择与混合时间对双时标SA收敛速度的影响。

提出的方法

  • 使用李雅普诺夫函数分析双时标系统中快速与慢速迭代之间的耦合。
  • 引入底层马尔可夫过程的几何混合时间,以量化并控制由依赖样本引入的偏差。
  • 采用时间尺度分离,其中 $\beta_k \ll \alpha_k$,$\alpha_k$ 与 $\beta_k$ 分别为快速与慢速迭代的步长。
  • 引入修正的误差过程 $\hat{z}_k = \hat{x}_k + \hat{y}_k$,以追踪两个迭代相对于其不动点的联合偏差。
  • 通过常数 $B$、$\mu_F$、$\mu_G$ 与混合时间 $\tau(\alpha_k)$ 对误差项应用递推不等式与有界性,推导收敛界。
  • 利用 $w_k$ 的指数加权控制误差项的增长,并推导期望平方误差的统一有界性。

实验结果

研究问题

  • RQ1当数据由马尔可夫过程生成时,非线性双时标随机逼近的有限时间收敛速率是什么?
  • RQ2马尔可夫过程的几何混合时间如何影响算法的偏差与收敛性?
  • RQ3奇异摄动框架能否扩展至马尔可夫采样下的非渐近收敛速率?
  • RQ4何种步长选择可确保在处理依赖观测时的最优收敛速度?
  • RQ5快速与慢速迭代之间的耦合,以及非独立同分布数据带来的偏差,如何共同影响均方误差?

主要发现

  • 该方法在迭代均方误差的期望下实现了 ${\cal O}(1/k^{2/3})$ 的有限时间收敛速率。
  • 该收敛速率在算子与马尔可夫过程的标准假设下推导得出,包括几何遍历性。
  • 马尔可夫链的几何混合时间被显式用于控制由依赖样本引入的偏差。
  • 分析表明,误差界依赖于步长乘积 $\alpha_k\beta_k$、$\alpha_k\alpha_{k;\tau(\alpha_k)}$ 与 $\beta_k^2$,其系数取决于算子的利普希茨常数与不动点范数。
  • 李雅普诺夫函数方法成功捕捉了两个时间尺度之间的相互作用,确保了稳定性与收敛性。
  • 该结果填补了文献中的关键空白,首次为马尔可夫噪声下的非线性双时标SA提供了有限时间收敛速率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。