Skip to main content
QUICK REVIEW

[Paper Review] Unregularized Online Learning Algorithms with General Loss Functions

Yiming Ying, Ding‐Xuan Zhou|arXiv (Cornell University)|Mar 2, 2015
Sparse and Compressive Sensing Techniques40 references3 citations
TL;DR

This paper establishes convergence rates for unregularized online learning algorithms in Reproducing Kernel Hilbert Spaces (RKHS) with general $α$-activating loss functions, using a novel induction-based proof technique. It provides explicit convergence rates under polynomially decaying step sizes and proves the convergence of the last iterate for both classification and pairwise learning tasks, extending prior results beyond regularized or Lipschitz-continuous gradient settings.

ABSTRACT

In this paper, we consider unregularized online learning algorithms in a Reproducing Kernel Hilbert Spaces (RKHS). Firstly, we derive explicit convergence rates of the unregularized online learning algorithms for classification associated with a general gamma-activating loss (see Definition 1 in the paper). Our results extend and refine the results in Ying and Pontil (2008) for the least-square loss and the recent result in Bach and Moulines (2011) for the loss function with a Lipschitz-continuous gradient. Moreover, we establish a very general condition on the step sizes which guarantees the convergence of the last iterate of such algorithms. Secondly, we establish, for the first time, the convergence of the unregularized pairwise learning algorithm with a general loss function and derive explicit rates under the assumption of polynomially decaying step sizes. Concrete examples are used to illustrate our main results. The main techniques are tools from convex analysis, refined inequalities of Gaussian averages, and an induction approach.

Motivation & Objective

  • To establish explicit convergence rates for unregularized online learning algorithms in RKHS with general $α$-activating loss functions.
  • To derive a general condition on step sizes ensuring convergence of the last iterate, extending beyond fixed or special forms of decay.
  • To extend theoretical results to pairwise learning, proving convergence and deriving rates for unregularized algorithms with general loss functions.
  • To overcome limitations of prior work by avoiding regularization and handling non-Lipschitz loss functions through refined convex analysis tools.

Proposed method

  • Uses a novel induction approach to analyze the convergence of the last iterate, avoiding reliance on regularization.
  • Applies refined inequalities of Gaussian and Rademacher averages to control the complexity of the hypothesis space.
  • Employs tools from convex analysis to handle the general $α$-activating loss function, defined by convexity, differentiability, and Hölder continuity of the derivative.
  • Derives recursive inequalities involving the error $R_t$, the step size $γ_t$, and norms of the kernel functions to bound the convergence rate.
  • Introduces a key recursive bound $R_{t+1} \leq F(R_t) + \text{error terms}$, where $F(R_t)$ captures the main decay behavior.
  • Establishes convergence by showing $R_t \leq D t^{-\beta}$ through induction, using a carefully chosen $D$ to satisfy the recursive inequality.

Experimental results

Research questions

  • RQ1Can explicit convergence rates be derived for unregularized online learning algorithms with general $α$-activating loss functions?
  • RQ2What general condition on step sizes ensures convergence of the last iterate in unregularized online learning?
  • RQ3Can the convergence of unregularized pairwise learning algorithms be established for general loss functions?
  • RQ4How do the convergence rates compare to existing results for regularized or least-square loss settings?
  • RQ5Can the assumption that the optimal function exists in the RKHS be removed for general loss functions?

Key findings

  • The paper establishes a convergence rate of $\mathcal{O}(T^{-1/3})$ for unregularized online learning with general $α$-activating loss functions under polynomially decaying step sizes.
  • A general condition on step sizes is derived that guarantees convergence of the last iterate, extending beyond the special case of $\mathcal{O}(t^{-\theta})$ decay.
  • For the first time, the paper proves convergence and derives explicit rates for unregularized pairwise learning algorithms with general loss functions.
  • The proof technique is shown to be simpler and more powerful than prior approaches, particularly for non-Lipschitz loss functions.
  • The results are established under a general framework that allows for $\alpha$-activating losses, including the least-square, logistic, and $q$-norm losses.
  • The analysis relies on refined inequalities of Gaussian averages and an induction-based argument to control the error propagation over iterations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.