Skip to main content
QUICK REVIEW

[论文解读] On Linear Stochastic Approximation: Fine-grained Polyak-Ruppert and Non-Asymptotic Concentration

Wenlong Mou, Chris Junchi Li|arXiv (Cornell University)|Apr 9, 2020
Stochastic Gradient Optimization Techniques参考文献 43被引用 21
一句话总结

本文对常步长下带Polyak-Ruppert平均的线性随机近似进行了精细化分析,建立了包含与步长相关的校正项的精确渐近协方差的中心极限定理。进一步推导出与CLT方差匹配至常数因子的非渐近浓度不等式,并证明了即使系统矩阵非Hurwitz但其特征值实部非负时,仍可实现O(1/T)的均方误差收敛速率。

ABSTRACT

We undertake a precise study of the asymptotic and non-asymptotic properties of stochastic approximation procedures with Polyak-Ruppert averaging for solving a linear system $\bar{A} θ= \bar{b}$. When the matrix $\bar{A}$ is Hurwitz, we prove a central limit theorem (CLT) for the averaged iterates with fixed step size and number of iterations going to infinity. The CLT characterizes the exact asymptotic covariance matrix, which is the sum of the classical Polyak-Ruppert covariance and a correction term that scales with the step size. Under assumptions on the tail of the noise distribution, we prove a non-asymptotic concentration inequality whose main term matches the covariance in CLT in any direction, up to universal constants. When the matrix $\bar{A}$ is not Hurwitz but only has non-negative real parts in its eigenvalues, we prove that the averaged LSA procedure actually achieves an $O(1/T)$ rate in mean-squared error. Our results provide a more refined understanding of linear stochastic approximation in both the asymptotic and non-asymptotic settings. We also show various applications of the main results, including the study of momentum-based stochastic gradient methods as well as temporal difference algorithms in reinforcement learning.

研究动机与目标

  • 为常步长线性随机近似中平均迭代序列的渐近分布提供精确表征。
  • 通过推导与CLT方差匹配的高概率浓度不等式,弥合渐近CLT结果与非渐近有限样本界之间的差距。
  • 将Polyak-Ruppert平均的理论理解从稳定(Hurwitz)系统扩展至特征值实部非负的情形。
  • 将精细化分析应用于基于动量的SGD与时序差分学习,揭示收敛性与误差行为的新见解。

提出的方法

  • 为常步长线性随机近似中平均迭代序列推导中心极限定理(CLT),其渐近协方差矩阵包含与步长成比例的校正项。
  • 利用迭代动态的鞅表示法,处理观测矩阵A与向量b中依赖于状态的噪声。
  • 证明迭代过程{θₜ}作为马氏链的遍历性,以支持Lindeberg型CLT与遍历定理的应用。
  • 应用伸缩恒等式以有界误差的二阶矩,实现非渐近浓度分析。
  • 对噪声分布施加更强的尾部假设,以推导出主导项与CLT方差匹配的高概率界。
  • 利用谱分解与通过特征向量矩阵U控制条件数,以管理噪声对当前迭代序列的依赖性。

实验结果

研究问题

  • RQ1在常步长线性随机近似中,平均迭代序列的精确渐近协方差是什么?它与经典Polyak-Ruppert形式有何不同?
  • RQ2能否为平均迭代序列推导出非渐近的高概率浓度界,使得主导项与渐近方差匹配?
  • RQ3当系统矩阵A的特征值实部非负(但不一定是Hurwitz)时,收敛速率如何变化?
  • RQ4线性随机近似的精细化结果如何应用于强化学习中的基于动量的SGD与时序差分学习?
  • RQ5该分析能否捕捉动量的加速效应,并提供近似最优速率的实例相关ℓ∞界?

主要发现

  • 为常步长Polyak-Ruppert平均在线性随机近似中建立了中心极限定理,其渐近协方差矩阵为经典Polyak-Ruppert项与与步长线性相关的校正项之和。
  • 在次高斯噪声假设下,推导出非渐近浓度不等式,其中任意方向的主导项与渐近方差匹配至绝对常数因子,且对失败概率具有对数多项式依赖。
  • 当矩阵Â的特征值实部非负(但不一定是Hurwitz)时,平均迭代序列在均方误差下实现O(1/T)收敛速率,且与谱间隙无关。
  • 通过一个新颖引理控制误差二阶矩的增长,该引理利用特征向量矩阵U的条件数与步长约束。
  • 分析为强化学习中策略评估提供了实例相关ℓ∞界,达到近似最优速率,且为平均奖励TD学习提供了与间隙无关的结果。
  • 结果解释了SGD中动量的加速效应,并为非线性设置下的线性化随机近似提供了精细化的理论理解。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。