Skip to main content
QUICK REVIEW

[Paper Review] Fast rates in statistical and online learning

Tim van Erven, Peter Grünwald|arXiv (Cornell University)|Jul 9, 2015
Machine Learning and Algorithms52 references4 citations
TL;DR

This paper unifies conditions for fast convergence rates in statistical and online learning by introducing the central condition for proper learning and stochastic mixability for online algorithms. It shows these conditions are equivalent under weak assumptions and generalize key concepts like the Tsybakov margin and Bernstein conditions, enabling $O(1/n)$ rates even for unbounded losses.

ABSTRACT

The speed with which a learning algorithm converges as it is presented with more data is a central problem in machine learning --- a fast rate of convergence means less data is needed for the same level of performance. The pursuit of fast rates in online and statistical learning has led to the discovery of many conditions in learning theory under which fast learning is possible. We show that most of these conditions are special cases of a single, unifying condition, that comes in two forms: the central condition for 'proper' learning algorithms that always output a hypothesis in the given model, and stochastic mixability for online algorithms that may make predictions outside of the model. We show that under surprisingly weak assumptions both conditions are, in a certain sense, equivalent. The central condition has a re-interpretation in terms of convexity of a set of pseudoprobabilities, linking it to density estimation under misspecification. For bounded losses, we show how the central condition enables a direct proof of fast rates and we prove its equivalence to the Bernstein condition, itself a generalization of the Tsybakov margin condition, both of which have played a central role in obtaining fast rates in statistical learning. Yet, while the Bernstein condition is two-sided, the central condition is one-sided, making it more suitable to deal with unbounded losses. In its stochastic mixability form, our condition generalizes both a stochastic exp-concavity condition identified by Juditsky, Rigollet and Tsybakov and Vovk's notion of mixability. Our unifying conditions thus provide a substantial step towards a characterization of fast rates in statistical learning, similar to how classical mixability characterizes constant regret in the sequential prediction with expert advice setting.

Motivation & Objective

  • To identify a single, unifying condition that explains fast convergence rates in both statistical and online learning settings.
  • To bridge the gap between proper learning (outputting hypotheses within the model) and online learning (allowing predictions outside the model) by introducing two forms of a core condition.
  • To show that the central condition and stochastic mixability are equivalent under weak assumptions, providing a unified framework for fast rates.
  • To demonstrate that the central condition generalizes the Bernstein condition and Tsybakov margin condition, especially for unbounded losses.
  • To provide a direct proof of fast rates under the central condition and link it to convexity of pseudoprobability sets under model misspecification.

Proposed method

  • Introduces the central condition for proper learning algorithms that always output a hypothesis in the model $\mathcal{F}$, ensuring fast $O(1/n)$ convergence.
  • Proposes stochastic mixability as the online counterpart, allowing algorithms to make predictions outside $\mathcal{F}$ while maintaining fast rates.
  • Establishes equivalence between the central condition and stochastic mixability under mild regularity assumptions.
  • Uses exponential moment bounds and the dominated convergence theorem to derive concentration inequalities for excess risk.
  • Applies union bounds over hypothesis classes with bounded metric entropy to control the probability of selecting suboptimal hypotheses.
  • Leverages the relationship between the central condition and convexity of pseudoprobability sets to connect to density estimation under misspecification.

Experimental results

Research questions

  • RQ1What single condition unifies fast convergence rates across statistical and online learning settings?
  • RQ2How do the central condition and stochastic mixability relate, and under what assumptions are they equivalent?
  • RQ3Can the central condition generalize the Bernstein and Tsybakov margin conditions, especially for unbounded losses?
  • RQ4What is the role of convexity in the set of pseudoprobabilities under the central condition?
  • RQ5How can fast $O(1/n)$ rates be achieved in the agnostic (non-realizable) setting using the central condition?

Key findings

  • The central condition enables a direct proof of $O(1/n)$ convergence rates for proper learning algorithms under surprisingly weak assumptions.
  • The central condition is one-sided, making it more suitable for unbounded losses than two-sided conditions like the Bernstein condition.
  • The central condition is equivalent to the Bernstein condition for bounded losses, and both generalize the Tsybakov margin condition.
  • Stochastic mixability generalizes both Vovk’s notion of mixability and the stochastic exp-concavity condition of Juditsky, Rigollet, and Tsybakov.
  • The central condition implies that the set of pseudoprobabilities is convex, linking it to density estimation under model misspecification.
  • For ERM, with high probability $1 - \delta$, the excess risk is bounded by $\frac{5\max\{V, 1/\eta^*\}(\log(1/\delta) + \log N)}{n}$, achieving fast rates under the central condition.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.