Skip to main content
QUICK REVIEW

[Paper Review] Sequential convergence of AdaGrad algorithm for smooth convex optimization

Cheik Traoré, Edouard Pauwels|arXiv (Cornell University)|Nov 24, 2020
Sparse and Compressive Sensing Techniques25 references32 citations
TL;DR

This paper proves that AdaGrad in both the scalar-step and coordinatewise variants produces convergent iterates when applied to convex functions with Lipschitz gradients, by establishing variable-metric quasi-Fejér monotonicity. The results show convergence to a global minimum without requiring a bounded domain.

ABSTRACT

We prove that the iterates produced by, either the scalar step size variant, or the coordinatewise variant of AdaGrad algorithm, are convergent sequences when applied to convex objective functions with Lipschitz gradient. The key insight is to remark that such AdaGrad sequences satisfy a variable metric quasi-Fej\\'er monotonicity property, which allows to prove convergence.

Motivation & Objective

  • Motivate the study of convergence of adaptive gradient methods in convex optimization.
  • Establish that AdaGrad variants produce convergent iterates to a global minimum under Lipschitz-gradient and minimum-attainment assumptions.
  • Introduce and utilize variable metric quasi-Fejér monotonicity to prove convergence.

Proposed method

  • Analyze two AdaGrad variants: AdaGrad-Norm with a scalar step size and AdaGrad with coordinatewise updates.
  • Show that both sequences are bounded and satisfy a variable-m metric quasi-Fejér monotonicity property relative to the set of minimizers.
  • Leverage a Descent Lemma for L-Lipschitz gradients and accumulate gradient norm bounds to prove convergence.
  • Prove summability of gradient norms, leading to convergence of the iterates to a minimizer.

Experimental results

Research questions

  • RQ1Do AdaGrad-Norm and AdaGrad with coordinatewise updates yield convergent sequences when F has a Lipschitz gradient and attains its minimum?
  • RQ2Can variable metric quasi-Fejér monotonicity be used to establish iterate convergence for adaptive gradient methods without a bounded domain assumption?

Key findings

  • AdaGrad-Norm and AdaGrad generate convergent sequences that converge to a global minimizer of F.
  • The gradient norms are summable under the given assumptions, implying convergence of the iterates.
  • The analysis does not require a bounded domain, unlike some prior results.
  • AdaGrad’s coordinatewise variant also converges under the same framework.
  • The convergence is established via variable metric quasi-Fejér monotonicity and related Lyapunov-like control.
  • The results hold for convex functions with Lipschitz gradients and a guaranteed minimum attainment.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.