Skip to main content
QUICK REVIEW

[Paper Review] Online estimation of the asymptotic variance for averaged stochastic gradient algorithms

Antoine Godichon‐Baggioni|arXiv (Cornell University)|Feb 3, 2017
Stochastic Gradient Optimization Techniques43 references15 citations
TL;DR

This paper proposes a recursive online algorithm for estimating the asymptotic variance of averaged stochastic gradient estimators in general Hilbert spaces, establishing almost sure and quadratic mean convergence rates. It extends asymptotic normality results for both standard and averaged stochastic gradient algorithms, enabling valid inference in high-dimensional and online learning settings.

ABSTRACT

Stochastic gradient algorithms are more and more studied since they can deal efficiently and online with large samples in high dimensional spaces. In this paper, we first establish a Central Limit Theorem for these estimates as well as for their averaged version in general Hilbert spaces. Moreover, since having the asymptotic normality of estimates is often unusable without an estimation of the asymptotic variance, we introduce a new recursive algorithm for estimating this last one, and we establish its almost sure rate of convergence as well as its rate of convergence in quadratic mean. Finally, two examples consisting in estimating the parameters of the logistic regression and estimating geometric quantiles are given.

Motivation & Objective

  • Address the lack of efficient, recursive asymptotic variance estimation for averaged stochastic gradient algorithms in high-dimensional and online settings.
  • Establish asymptotic normality for both standard and averaged stochastic gradient estimators in general Hilbert spaces.
  • Provide theoretical convergence rates for the proposed recursive variance estimator under mild regularity and local strong convexity assumptions.
  • Enable practical statistical inference (e.g., confidence intervals) for stochastic gradient methods in large-scale and sequential data contexts.
  • Demonstrate applicability through logistic regression and geometric quantile estimation in high-dimensional and functional data settings.

Proposed method

  • Derive a Central Limit Theorem for averaged stochastic gradient estimators in separable Hilbert spaces under local strong convexity and moment conditions.
  • Propose a recursive algorithm inspired by Gahbiche and Pelletier (2000) to estimate the asymptotic variance of the averaged estimator.
  • Establish almost sure convergence and $ L^2 $ convergence rates for the recursive variance estimator using martingale difference sequences and moment bounds.
  • Utilize a functional framework with Hilbert-Schmidt operators to handle general Hilbert space-valued gradients and their covariance structure.
  • Apply technical lemmas on moment bounds and exponential sums to control the bias and variance of the recursive estimator.
  • Ensure recursive computation by updating the variance estimate incrementally using online observations, avoiding storage of all past data.

Experimental results

Research questions

  • RQ1Can asymptotic normality be established for averaged stochastic gradient estimators in general Hilbert spaces under local strong convexity?
  • RQ2What recursive algorithm enables consistent and convergent online estimation of the asymptotic variance of averaged stochastic gradient estimators?
  • RQ3What are the almost sure and quadratic mean convergence rates of the proposed recursive variance estimator?
  • RQ4How does the proposed method enable valid statistical inference (e.g., confidence intervals) in high-dimensional and online learning settings?
  • RQ5Can the method be effectively applied to real-world problems such as logistic regression and geometric quantile estimation?

Key findings

  • The paper establishes a Central Limit Theorem for averaged stochastic gradient estimators in separable Hilbert spaces under mild regularity and local strong convexity assumptions.
  • The proposed recursive algorithm for asymptotic variance estimation achieves almost sure convergence and $ L^2 $ convergence at a rate that depends on the step-size schedule and problem geometry.
  • The convergence rates are formally derived using moment bounds and exponential sum inequalities, with explicit dependence on the distance to the optimal solution.
  • The method is applicable to both parametric models (e.g., logistic regression) and nonparametric robust statistics (e.g., geometric quantiles).
  • The recursive variance estimator is shown to be consistent and computationally efficient, requiring only online updates without storing past data.
  • Simulation results and theoretical analysis confirm the method’s effectiveness in high-dimensional and functional data settings with sequential data streams.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.