[Paper Review] Stochastic subgradient method converges at the rate $O(k^{-1/4})$ on weakly convex functions
The paper shows that the proximal stochastic subgradient method applied to weakly convex objectives drives the gradient of the Moreau envelope to zero at rate O(k^{-1/4}), yielding an O(ε^{-4}) iteration complexity to obtain near-stationarity.
We prove that the proximal stochastic subgradient method, applied to a weakly convex problem, drives the gradient of the Moreau envelope to zero at the rate $O(k^{-1/4})$. As a consequence, we resolve an open question on the convergence rate of the proximal stochastic gradient method for minimizing the sum of a smooth nonconvex function and a convex proximable function.
Motivation & Objective
- Motivate and analyze optimization of φ(x) = g(x) + r(x) where r is convex proximable and g is ρ-weakly convex.
- Provide convergence guarantees for the proximal stochastic subgradient method under standard stochastic oracle assumptions (A1–A3).
- Characterize the rate of near-stationarity via the gradient of the Moreau envelope φ_{λ}, with λ = 1/(2ρ).
- Show that the method achieves an ε-stationarity measure in O(ε^{-4}) iterations (for appropriate settings).
- Discuss how these results extend known rates to nonsmooth g and allow non-decreasing variance in the stochastic estimates.
Proposed method
- Formulate the problem with φ(x)=g(x)+r(x) where r is closed convex with computable proximal map and g is ρ-weakly convex.
- Use proximal stochastic subgradient updates x_{t+1} = prox_{α_t r}(x_t - α_t G(x_t, ξ_t)) with G(x_t, ξ_t) an unbiased estimator of a subgradient of g.
- Define Moreau envelope φ_λ and use the gradient ∇φ_{λ}(x) = (x - prox_{λφ}(x))/λ to measure near-stationarity.
- Prove convergence under assumptions (A1) i.i.d. data, (A2) stochastic subgradient in ∂g(x), (A3) bounded variance of G, and α_t in (0, 1/ρ].
- Derive bounds on E[||∇φ_{1/ârho}(x_{t*})||^2] in terms of initial gap, variance, and step sizes; obtain O(ε^{-4}) iteration complexity for ε-stationarity.
- Provide corollaries for constant stepsizes and discuss improvements in the convex/smooth cases.
Experimental results
Research questions
- RQ1What is the convergence rate of the proximal stochastic subgradient method for weakly convex objectives?
- RQ2Can near-stationarity be certified via the gradient of the Moreau envelope, and at what rate does ∥∇φ_{1/(2ρ)}(x)∥ shrink?
- RQ3How do variance assumptions of the stochastic oracle affect the rate, and can non-diminishing variance be tolerated?
- RQ4What are the iteration complexities to achieve ε-stationarity under the proposed framework?
- RQ5How do results adapt when g is smooth or r is an indicator/projection term?
Key findings
- The proximal stochastic subgradient method drives the gradient of the Moreau envelope to zero at rate O(k^{-1/4}).
- Under standard assumptions, the method yields E[∥∇φ_{1/(2ρ)}(x_{t*})∥^2] ≤ C/(√{T+1}) with appropriate constants, implying ε-stationarity in O(ε^{-4}) iterations.
- With constant step size α ≈ 1/√(T+1), the bound scales as O( (φ_{1/(2ρ)}(x0) - min φ) + ρ L^2 γ^2 ) / (γ √(T+1)).
- If g is convex, the paper outlines potential improvements via multi-stage or regularized variants, achieving faster rates in certain regimes.
- In the smooth setting with finite variance, analogous bounds hold for ∥∇φ_{1/(2ρ)}(x_{t*})∥^2, with a similar ε^{-4} dependence and additional σ^2 terms.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.