Skip to main content
QUICK REVIEW

[Paper Review] On Empirical Risk Minimization with Dependent and Heavy-Tailed Data

Abhishek Roy, Krishnakumar Balasubramanian|arXiv (Cornell University)|Sep 6, 2021
Statistical Methods and Inference32 references4 citations
TL;DR

This paper extends empirical risk minimization (ERM) to dependent and heavy-tailed data by generalizing Mendelson's learning without concentration framework to strictly stationary, exponentially β-mixing processes. It establishes risk bounds under weaker moment assumptions by directly modeling the interaction between noise and function evaluations, deriving convergence rates for high-dimensional linear regression under squared and Huber losses with polynomially or exponentially heavy-tailed errors.

ABSTRACT

In this work, we establish risk bounds for the Empirical Risk Minimization (ERM) with both dependent and heavy-tailed data-generating processes. We do so by extending the seminal works of Mendelson [Men15, Men18] on the analysis of ERM with heavy-tailed but independent and identically distributed observations, to the strictly stationary exponentially $β$-mixing case. Our analysis is based on explicitly controlling the multiplier process arising from the interaction between the noise and the function evaluations on inputs. It allows for the interaction to be even polynomially heavy-tailed, which covers a significantly large class of heavy-tailed models beyond what is analyzed in the learning theory literature. We illustrate our results by deriving rates of convergence for the high-dimensional linear regression problem with dependent and heavy-tailed data.

Motivation & Objective

  • To close the theoretical gap in ERM analysis for dependent and heavy-tailed data-generating processes (DGPs), which are common in practice but poorly understood in learning theory.
  • To extend Mendelson's learning without concentration framework—originally for i.i.d. heavy-tailed data—to the non-i.i.d. case with strictly stationary, exponentially β-mixing DGPs.
  • To derive risk bounds for convex, locally strongly-convex loss functions when the noise and input interaction may be polynomially or exponentially heavy-tailed.
  • To provide convergence rates for high-dimensional linear regression under both squared and Huber losses with dependent, heavy-tailed data.
  • To develop new concentration inequalities for β-mixing random variables to handle multiplier processes under weak moment conditions.

Proposed method

  • Extends the small-ball technique from Mendelson (2015, 2018) to β-mixing processes by directly assuming moment conditions on the interaction between noise and function evaluations.
  • Uses concentration inequalities from [MPR11] for exponentially heavy-tailed interactions and extends [BMdlP20] to derive new inequalities for polynomially heavy-tailed interactions in the β-mixing setting.
  • Analyzes the multiplier empirical process by explicitly controlling the interaction term between noise and function evaluations on inputs, avoiding reliance on uniform concentration.
  • Applies the framework to high-dimensional linear models with function classes defined by ℓ1- and ℓ2-balls, using complexity measures like ωQ(F−F,N,ζ1,ζ2) to bound the estimation error.
  • Derives generalization error bounds via Corollary B.2, combining tail probability bounds with complexity control under β-mixing assumptions.
  • Establishes convergence rates by balancing three terms in the risk bound: a fast rate term, a slow rate term, and a term decaying with N via careful choice of the set A(N).

Experimental results

Research questions

  • RQ1Can ERM be theoretically justified for dependent, heavy-tailed data when the i.i.d. assumption fails?
  • RQ2How can multiplier empirical process inequalities be extended to β-mixing processes under weak moment assumptions?
  • RQ3What convergence rates can be achieved for high-dimensional linear regression under squared and Huber losses when the data are dependent and heavy-tailed?
  • RQ4Can the learning without concentration framework be adapted to non-i.i.d. settings without relying on uniform concentration or boundedness assumptions?
  • RQ5What are the implications of polynomially versus exponentially heavy-tailed interactions in the noise-function evaluation product for generalization error?

Key findings

  • The paper establishes a general risk bound for ERM under β-mixing, heavy-tailed data, with convergence rates of order max(N^{-1/2+ι}, R/√N √log(ed/N)) for the ℓ2 estimation error.
  • For the high-dimensional linear model with ℓ1-regularized function class, the estimation error is bounded by O(R/√N √log(ed/N)) when N ≤ c1(Lg)d, and O(R/√d) when N > c1(Lg)d.
  • The generalization error bound holds with high probability, where the failure probability decays as exp(−cN^{η1/(1+η1)}) and exp(−(N^{2ι}τ0)^η/M1), indicating strong concentration.
  • The method achieves non-asymptotic risk bounds under weaker moment assumptions than classical ERM, allowing for polynomially or exponentially heavy-tailed interactions between noise and inputs.
  • The analysis shows that the third term in the risk bound (arising from the β-mixing dependence) can be made to decay to zero as N→∞ by choosing A(N) appropriately, enabling fast rates.
  • The results are validated in the context of sparse linear regression with both squared and Huber loss, demonstrating robustness to heavy-tailed noise under dependence.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.