Skip to main content
QUICK REVIEW

[论文解读] Convergence Analysis of Riemannian Stochastic Approximation Schemes

Alain Durmus, Pablo Jiménez|arXiv (Cornell University)|May 27, 2020
Stochastic Gradient Optimization Techniques参考文献 31被引用 4
一句话总结

本文针对流形上的黎曼随机近似方案提出了全局收敛性分析,允许存在有偏向量场、受控马尔可夫采样以及使用收缩映射替代指数映射。即使在不假设先验有界性或测地线凸性的条件下,也建立了期望平方范数的均场收敛速率 𝒪(b∞ + log n / √n)。

ABSTRACT

This paper analyzes the convergence for a large class of Riemannian stochastic approximation (SA) schemes, which aim at tackling stochastic optimization problems. In particular, the recursions we study use either the exponential map of the considered manifold (geodesic schemes) or more general retraction functions (retraction schemes) used as a proxy for the exponential map. Such approximations are of great interest since they are low complexity alternatives to geodesic schemes. Under the assumption that the mean field of the SA is correlated with the gradient of a smooth Lyapunov function (possibly non-convex), we show that the above Riemannian SA schemes find an ${\mathcal{O}}(b_\infty + \log n / \sqrt{n})$-stationary point (in expectation) within ${\mathcal{O}}(n)$ iterations, where $b_\infty \geq 0$ is the asymptotic bias. Compared to previous works, the conditions we derive are considerably milder. First, all our analysis are global as we do not assume iterates to be a-priori bounded. Second, we study biased SA schemes. To be more specific, we consider the case where the mean-field function can only be estimated up to a small bias, and/or the case in which the samples are drawn from a controlled Markov chain. Third, the conditions on retractions required to ensure convergence of the related SA schemes are weak and hold for well-known examples. We illustrate our results on three machine learning problems.

研究动机与目标

  • 在较弱假设下,分析一般黎曼流形上黎曼随机近似方案的收敛性。
  • 将收敛保证的适用范围从标准的指数映射扩展到计算上更可行的收缩映射方案。
  • 在随机近似框架中处理有偏均场估计器和马尔可夫链采样。
  • 消除对迭代序列先验有界性及测地线凸性的依赖,从而实现全局收敛结果。
  • 在温和的正则性条件下,为测地线和基于收缩映射的方案提供非渐近收敛速率。

提出的方法

  • 将流形上的根求解问题表述为寻找满足 h(θ) = 0 的 θ,其中 h 是通过在状态空间上积分定义的均向量场。
  • 提出一种使用指数映射或收缩映射更新迭代的黎曼随机近似方案:θ_{n+1} = Retr_{θ_n}(η_{n+1} H_{θ_n}(X_{n+1}))。
  • 在弱假设下分析收敛性:无需迭代序列的先验有界性,无需测地线凸性,且允许存在有偏或马尔可夫链采样。
  • 应用适配于黎曼几何的 ODE 方法,研究随机过程的极限行为。
  • 利用黎曼几何工具,包括平行移动、Hessian 有界性和曲率控制,推导稳定性和收敛性性质。
  • 通过局部鞅技术与矩界分析建立收敛速率,利用流形结构及收缩映射的性质。

实验结果

研究问题

  • RQ1在不假设迭代序列有界或测地线凸性的情况下,黎曼随机近似能否实现全局收敛?
  • RQ2向量场估计器中的偏差如何影响黎曼 SA 方案的收敛速率?
  • RQ3当采样分布通过受控马尔可夫链依赖于当前迭代时,收敛保证如何成立?
  • RQ4在黎曼 SA 中,收缩映射在多大程度上可以替代指数映射,同时保持收敛性?
  • RQ5在对均场和采样机制的假设较弱时,黎曼 SA 的非渐近收敛速率可如何推导?

主要发现

  • 所提出的黎曼随机近似方案对均场期望平方范数的收敛速率为 𝒪(b∞ + log n / √n),其中 b∞ 为渐近偏差。
  • 在无需迭代序列先验有界或流形紧致的条件下,建立了全局收敛性。
  • 分析在收缩映射的弱假设下成立,这些假设在许多实际场景中均满足,包括基于 QR 分解或极坐标收缩等常见形式。
  • 收敛结果可扩展至样本由受控马尔可夫链生成的情形,从而拓宽了其在在线学习和强化学习场景中的适用性。
  • 该框架适用于非凸问题,如主成分分析和黎曼中位数计算,理论保证收敛至临界点。
  • 本文为在黎曼优化中使用计算成本更低的收缩映射替代指数映射提供了理论基础,且不损失收敛保证。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。