Skip to main content
QUICK REVIEW

[Paper Review] High-Probability Bounds for Stochastic Optimization and Variational Inequalities: the Case of Unbounded Variance

Abdurakhmon Sadiev, М. А. Данилова|arXiv (Cornell University)|Feb 2, 2023
Sparse and Compressive Sensing Techniques5 citations
TL;DR

This paper presents high-probability convergence guarantees for stochastic optimization and variational inequalities under relaxed assumptions, specifically allowing unbounded variance by assuming only bounded central $α$-th moments of gradient and operator noise for $\alpha \in (1,2]$. The authors develop novel algorithms with improved robustness, achieving $\widetilde{\cal O}(\cdot)$ complexity bounds in smooth non-convex, convex, and monotone settings without requiring bounded gradients or variance.

ABSTRACT

During recent years the interest of optimization and machine learning communities in high-probability convergence of stochastic optimization methods has been growing. One of the main reasons for this is that high-probability complexity bounds are more accurate and less studied than in-expectation ones. However, SOTA high-probability non-asymptotic convergence results are derived under strong assumptions such as the boundedness of the gradient noise variance or of the objective's gradient itself. In this paper, we propose several algorithms with high-probability convergence results under less restrictive assumptions. In particular, we derive new high-probability convergence results under the assumption that the gradient/operator noise has bounded central $α$-th moment for $α\in (1,2]$ in the following setups: (i) smooth non-convex / Polyak-Lojasiewicz / convex / strongly convex / quasi-strongly convex minimization problems, (ii) Lipschitz / star-cocoercive and monotone / quasi-strongly monotone variational inequalities. These results justify the usage of the considered methods for solving problems that do not fit standard functional classes studied in stochastic optimization.

Motivation & Objective

  • To address the gap in high-probability convergence analysis for stochastic optimization under weak noise assumptions.
  • To extend existing results beyond bounded variance or gradient assumptions, which restrict applicability to real-world problems with heavy-tailed noise.
  • To develop algorithms that maintain high-probability convergence in smooth non-convex, convex, and monotone variational inequality problems with unbounded variance.
  • To provide complexity bounds that are sensitive to noise tail behavior, improving practical relevance over in-expectation bounds.
  • To establish theoretical guarantees for gradient clipping and adaptive step-sizes under $\alpha$-moment conditions.

Proposed method

  • Proposes a novel analysis framework based on $\alpha$-th moment bounds ($\alpha \in (1,2]$) for gradient and operator noise, replacing the standard bounded variance assumption.
  • Introduces a modified stochastic gradient method with adaptive step-sizes and gradient clipping to control large deviations in heavy-tailed noise settings.
  • Employs a recursive concentration argument using Bernstein-type inequalities to bound the cumulative noise in iterates, ensuring high-probability stability.
  • Derives high-probability bounds for both minimization and variational inequality problems using restricted gap functions and level-set invariance.
  • Uses induction and tail probability control to show that iterates remain within a ball around the solution with high probability, even under unbounded variance.
  • Establishes complexity bounds via a trade-off between step-size choice and confidence level, balancing convergence rate and robustness.

Experimental results

Research questions

  • RQ1Can high-probability convergence be achieved for stochastic optimization when the gradient noise has unbounded variance?
  • RQ2What are the minimal assumptions on noise moments that still allow non-asymptotic high-probability convergence?
  • RQ3How do $\alpha$-moment conditions ($\alpha \in (1,2]$) affect the convergence rate and complexity of stochastic methods?
  • RQ4Can gradient clipping and adaptive step-sizes ensure high-probability convergence under heavy-tailed noise?
  • RQ5What is the optimal trade-off between confidence level $1-\beta$ and iteration complexity under $\alpha$-moment noise?

Key findings

  • The paper establishes high-probability convergence for smooth non-convex, convex, and strongly convex minimization under $\alpha$-moment bounded noise, with $\alpha \in (1,2]$.
  • For strongly convex problems, the method achieves $\|x^{K+1} - x^* \|^{2} = \widetilde{\cal O}\left(\max\left\{\exp\left(-\frac{\mu K}{\ell \ln \frac{K}{\beta}}\right), \frac{\sigma^2 \ln^{\frac{2(\alpha-1)}{\alpha}}(K/\beta) \ln^2(B_\varepsilon)}{K^{\frac{2(\alpha-1)}{\alpha}} \mu^2}\right\} \right)$ with probability at least $1-\beta$.
  • The iteration complexity for achieving $\|x^{K+1} - x^* \|^{2} \leq \varepsilon$ is $K = \widetilde{\cal O}\left(\frac{\ell}{\mu} \ln\left(\frac{R^2}{\varepsilon}\right) \ln\left(\frac{\ell}{\mu\beta} \ln\frac{R^2}{\varepsilon}\right) + \left(\frac{\sigma^2}{\mu^2 \varepsilon}\right)^{\frac{\alpha}{2(\alpha-1)}} \ln\left(\frac{1}{\beta} \left(\frac{\sigma^2}{\mu^2 \varepsilon}\right)^{\frac{\alpha}{2(\alpha-1)}}\right) \ln^{\frac{\alpha}{\alpha-1}}(B_\varepsilon)\right)$.
  • The analysis shows that the iterates remain within a ball of radius $R$ around the solution with high probability, even under unbounded variance, by controlling the growth of noise terms.
  • For monotone and strongly monotone variational inequalities, the method achieves $\widetilde{\cal O}(\cdot)$ complexity bounds under the same $\alpha$-moment assumption.
  • The results demonstrate that high-probability convergence is possible without assuming bounded gradients or variance, significantly broadening the scope of theoretical guarantees.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.