Skip to main content
QUICK REVIEW

[论文解读] On the Stability of Random Matrix Product with Markovian Noise: Application to Linear Stochastic Approximation and TD Learning

Alain Durmus, Éric Moulines|arXiv (Cornell University)|Jan 30, 2021
Reinforcement Learning in Robotics参考文献 19被引用 5
一句话总结

本论文在放宽条件的前提下,为由一般马尔可夫链驱动的随机矩阵乘积建立了指数稳定性:即超李雅普诺夫漂移条件与矩阵函数受控增长。该研究推导出在马尔可夫噪声下线性随机逼近与时序差分学习的有限时间p阶矩界,将先前结果扩展至无界状态空间和无界矩阵函数的情形。

ABSTRACT

This paper studies the exponential stability of random matrix products driven by a general (possibly unbounded) state space Markov chain. It is a cornerstone in the analysis of stochastic algorithms in machine learning (e.g. for parameter tracking in online learning or reinforcement learning). The existing results impose strong conditions such as uniform boundedness of the matrix-valued functions and uniform ergodicity of the Markov chains. Our main contribution is an exponential stability result for the $p$-th moment of random matrix product, provided that (i) the underlying Markov chain satisfies a super-Lyapunov drift condition, (ii) the growth of the matrix-valued functions is controlled by an appropriately defined function (related to the drift condition). Using this result, we give finite-time $p$-th moment bounds for constant and decreasing stepsize linear stochastic approximation schemes with Markovian noise on general state space. We illustrate these findings for linear value-function estimation in reinforcement learning. We provide finite-time $p$-th moment bound for various members of temporal difference (TD) family of algorithms.

研究动机与目标

  • 解决在一般马尔可夫噪声下线性随机逼近(LSA)与时序差分(TD)学习缺乏有限时间矩界的空白。
  • 克服先前研究中要求统一几何遍历性(UGE)与矩阵函数一致有界的局限性。
  • 将稳定性分析扩展至无界状态空间,并适用于潜在无界的矩阵值函数的LSA与TD学习。
  • 为在弱遍历性条件下分析随机矩阵乘积的p阶矩稳定性提供一个通用框架。
  • 使强化学习算法在实际无界场景中实现有限时间性能保证成为可能。

提出的方法

  • 引入随机矩阵乘积的(V,q)-指数稳定性条件,其中V是由超李雅普诺夫漂移条件导出的李雅普诺夫函数。
  • 通过与漂移条件相关的函数W₁和W₂定义矩阵增长界,允许矩阵无界。
  • 采用一种新颖的李雅普诺夫函数构造方法,结合时齐马尔可夫核与几何矩生成函数。
  • 通过耦合论证与小集的子水平集论证,建立随机矩阵乘积的漂移条件。
  • 将稳定性结果应用于将LSA中的误差分解为加权矩阵乘积与初始误差之和,从而实现矩界推导。
  • 采用时域耦合技术控制矩阵乘积的尾部行为,并通过指数矩不等式推导出矩界。

实验结果

研究问题

  • RQ1在不依赖统一几何遍历性的情况下,能否基于超李雅普诺夫漂移条件建立随机矩阵乘积的指数稳定性?
  • RQ2当矩阵函数无界时,如何为具有马尔可夫噪声的线性随机逼近推导出有限时间p阶矩界?
  • RQ3在一般(可能无界的)状态空间马尔可夫链下,矩阵函数增长的何种条件可确保矩稳定性?
  • RQ4所提出的框架能否应用于时序差分学习算法,以推导出有限时间性能保证?
  • RQ5李雅普诺夫函数与漂移条件在控制高阶矩中矩阵乘积增长方面起什么作用?

主要发现

  • 论文在超李雅普诺夫漂移条件与受控矩阵增长下,建立了随机矩阵乘积的(V,q)-指数稳定性,即使矩阵函数无界亦成立。
  • 为在一般状态空间上具有马尔可夫噪声的常数步长与递减步长线性随机逼近方案,推导出有限时间p阶矩界。
  • 与先前工作相比,该结果在更弱条件下成立:无需统一几何遍历性,也无需A(z)与b(z)有界。
  • 对于时序差分学习,该框架可为TD家族的多种成员(包括TD(0)、TD(λ)及带函数逼近的TD(λ))提供有限时间p阶矩界。
  • 分析中提出了一种新颖的李雅普诺夫函数构造方法,结合几何矩生成函数与子水平集论证,以证明小集条件。
  • 关键技术突破在于引入时域耦合与精心调节的参数β₀,以控制矩界中指数衰减率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。