[论文解读] High-Probability Bounds for Stochastic Optimization and Variational Inequalities: the Case of Unbounded Variance
本文在放宽假设条件下,为随机优化和变分不等式提供了高概率收敛保证,具体而言,允许方差无界,仅假设梯度和算子噪声的有界中心 $\alpha$-阶矩,其中 $\alpha \in (1,2]$。作者提出了新颖的算法,增强了鲁棒性,在光滑非凸、凸和单调设定下实现了 $\widetilde{\cal O}(\cdot)$ 复杂度界,且无需假设有界梯度或方差。
During recent years the interest of optimization and machine learning communities in high-probability convergence of stochastic optimization methods has been growing. One of the main reasons for this is that high-probability complexity bounds are more accurate and less studied than in-expectation ones. However, SOTA high-probability non-asymptotic convergence results are derived under strong assumptions such as the boundedness of the gradient noise variance or of the objective's gradient itself. In this paper, we propose several algorithms with high-probability convergence results under less restrictive assumptions. In particular, we derive new high-probability convergence results under the assumption that the gradient/operator noise has bounded central $α$-th moment for $α\in (1,2]$ in the following setups: (i) smooth non-convex / Polyak-Lojasiewicz / convex / strongly convex / quasi-strongly convex minimization problems, (ii) Lipschitz / star-cocoercive and monotone / quasi-strongly monotone variational inequalities. These results justify the usage of the considered methods for solving problems that do not fit standard functional classes studied in stochastic optimization.
研究动机与目标
- 解决在弱噪声假设下,随机优化中高概率收敛分析的空白。
- 将现有结果扩展至超越有界方差或梯度假设的范围,从而提升在具有重尾噪声的实际问题中的适用性。
- 开发在无界方差下,仍能保持高概率收敛的算法,适用于光滑非凸、凸及单调变分不等式问题。
- 提供对噪声尾部行为敏感的复杂度界,相较于期望值边界,提升实际相关性。
- 在 $\alpha$-矩条件下,建立对梯度裁剪和自适应步长的理论保证。
提出的方法
- 提出一种基于 $\alpha$-阶矩界($\alpha \in (1,2]$)的新型分析框架,用于梯度和算子噪声,替代标准的有界方差假设。
- 引入一种改进的随机梯度方法,结合自适应步长和梯度裁剪,以控制重尾噪声环境下的大偏差。
- 采用基于伯恩斯坦型不等式的递归集中论证,以界累积噪声在迭代中的影响,确保高概率稳定性。
- 利用受限间隙函数和水平集不变性,推导出最小化和变分不等式问题的高概率边界。
- 通过归纳法和尾部概率控制,证明即使在无界方差下,迭代点仍以高概率保持在解附近球体内。
- 通过步长选择与置信水平之间的权衡,建立复杂度界,平衡收敛速度与鲁棒性。
实验结果
研究问题
- RQ1当梯度噪声具有无界方差时,能否实现随机优化的高概率收敛?
- RQ2在仍能实现非渐近高概率收敛的前提下,噪声矩的最小假设是什么?
- RQ3$\alpha$-矩条件($\alpha \in (1,2]$)如何影响随机方法的收敛速度与复杂度?
- RQ4在重尾噪声下,梯度裁剪和自适应步长能否确保高概率收敛?
- RQ5在 $\alpha$-矩噪声下,置信水平 $1-\beta$ 与迭代复杂度之间的最优权衡是什么?
主要发现
- 本文在 $\alpha$-矩有界噪声($\alpha \in (1,2]$)下,建立了光滑非凸、凸及强凸最小化问题的高概率收敛性。
- 对于强凸问题,该方法以至少 $1-\beta$ 的概率实现 $\|x^{K+1} - x^*\|^{2} = \widetilde{\cal O}\left(\max\left\{\exp\left(-\frac{\mu K}{\ell \ln \frac{K}{\beta}}\right), \frac{\sigma^2 \ln^{\frac{2(\alpha-1)}{\alpha}}(K/\beta) \ln^2(B_\varepsilon)}{K^{\frac{2(\alpha-1)}{\alpha}} \mu^2}\right\} \right)$。
- 达到 $\|x^{K+1} - x^*\|^{2} \leq \varepsilon$ 所需的迭代复杂度为 $K = \widetilde{\cal O}\left(\frac{\ell}{\mu} \ln\left(\frac{R^2}{\varepsilon}\right) \ln\left(\frac{\ell}{\mu\beta} \ln\frac{R^2}{\varepsilon}\right) + \left(\frac{\sigma^2}{\mu^2 \varepsilon}\right)^{\frac{\alpha}{2(\alpha-1)}} \ln\left(\frac{1}{\beta} \left(\frac{\sigma^2}{\mu^2 \varepsilon}\right)^{\frac{\alpha}{2(\alpha-1)}}\right) \ln^{\frac{\alpha}{\alpha-1}}(B_\varepsilon)\right)$。
- 分析表明,通过控制噪声项的增长,即使在无界方差下,迭代点仍以高概率保持在解附近半径为 $R$ 的球体内。
- 对于单调和强单调变分不等式问题,该方法在相同的 $\alpha$-矩假设下,实现了 $\widetilde{\cal O}(\cdot)$ 复杂度界。
- 结果表明,无需假设有界梯度或方差,高概率收敛是可能的,显著拓宽了理论保证的适用范围。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。