Skip to main content
QUICK REVIEW

[Paper Review] Out-of-sample error estimate for robust M-estimators with convex penalty

Pierre Bellec|arXiv (Cornell University)|Aug 26, 2020
Statistical Methods and Inference40 references4 citations
TL;DR

This paper introduces a generic, data-driven out-of-sample error estimate for high-dimensional M-estimators regularized with convex penalties, applicable to robust losses like Huber and least squares. The estimate achieves n⁻¹/² relative error under mild regularity and sparsity conditions, enabling consistent noise variance and generalization error estimation without solving complex nonlinear systems or assuming isotropic design.

ABSTRACT

A generic out-of-sample error estimate is proposed for robust $M$-estimators regularized with a convex penalty in high-dimensional linear regression where $(X,y)$ is observed and $p,n$ are of the same order. If $ψ$ is the derivative of the robust data-fitting loss $ρ$, the estimate depends on the observed data only through the quantities $\hatψ= ψ(y-X\hatβ)$, $X^ op \hatψ$ and the derivatives $(\partial/\partial y) \hatψ$ and $(\partial/\partial y) X\hatβ$ for fixed $X$. The out-of-sample error estimate enjoys a relative error of order $n^{-1/2}$ in a linear model with Gaussian covariates and independent noise, either non-asymptotically when $p/n\le γ$ or asymptotically in the high-dimensional asymptotic regime $p/n oγ'\in(0,\infty)$. General differentiable loss functions $ρ$ are allowed provided that $ψ=ρ'$ is 1-Lipschitz. The validity of the out-of-sample error estimate holds either under a strong convexity assumption, or for the $\ell_1$-penalized Huber M-estimator if the number of corrupted observations and sparsity of the true $β$ are bounded from above by $s_*n$ for some small enough constant $s_*\in(0,1)$ independent of $n,p$. For the square loss and in the absence of corruption in the response, the results additionally yield $n^{-1/2}$-consistent estimates of the noise variance and of the generalization error. This generalizes, to arbitrary convex penalty, estimates that were previously known for the Lasso.

Motivation & Objective

  • To develop a generic, data-driven out-of-sample error estimate for M-estimators with convex penalties in high-dimensional linear regression.
  • To ensure the estimate is valid under minimal assumptions on the loss function and penalty, including robustness to heavy-tailed or contaminated responses.
  • To extend consistent noise variance and generalization error estimation beyond the Lasso to arbitrary convex penalties and general covariance structures.
  • To avoid reliance on isotropic design assumptions, allowing general covariance matrices Σ ≠ I.
  • To provide non-asymptotic guarantees that hold when p/n ≤ γ and asymptotically as p/n → γ′ ∈ (0, ∞).

Proposed method

  • Proposes a data-driven estimator of the out-of-sample error ∥Σ¹ᐟ²(bβ − β)∥² using only the residual derivative ψ(y − Xbβ), the quantity Σ⁻¹ᐟ²Xᵀψ(y − Xbβ), and the derivatives of bβ and ψ(y − Xbβ) with respect to y.
  • Derives explicit forms for the derivative of the estimator with respect to the design matrix X, using the chain rule and the KKT conditions of the optimization problem.
  • Applies concentration inequalities and random matrix theory to bound the error of the estimate under high-dimensional asymptotics.
  • Uses a perturbation argument and the Gaussian min-max theorem (GMT) framework to control the stability of the estimator under model perturbations.
  • Establishes the estimate’s validity under strong convexity or sparsity assumptions, with bounds on contamination and tuning parameters.
  • For the Huber Lasso, derives a closed-form expression for the estimate involving the active set size |Ŝ| and the norm of the projected residual.

Experimental results

Research questions

  • RQ1Can a generic, data-driven out-of-sample error estimate be constructed for M-estimators with convex penalties in high-dimensional settings?
  • RQ2Does the proposed estimate achieve n⁻¹/² relative error under general covariance and robust loss functions?
  • RQ3Can the estimate consistently estimate noise variance and generalization error in the absence of response contamination?
  • RQ4How does the method perform under non-isotropic design matrices, and can it avoid the isotropy assumption common in prior work?
  • RQ5What conditions on sparsity and tuning parameters ensure the validity of the error estimate for the Lasso and Huber Lasso?

Key findings

  • The proposed out-of-sample error estimate achieves a relative error of order n⁻¹/² in linear models with Gaussian covariates and independent noise, both non-asymptotically and asymptotically as p/n → γ′ ∈ (0, ∞).
  • The estimate is valid for general convex, differentiable, 1-Lipschitz-derivative loss functions, including least squares and Huber loss.
  • For the square loss and in the absence of response contamination, the method yields n⁻¹/²-consistent estimates of the noise variance and generalization error.
  • The method does not require solving systems of nonlinear equations or knowledge of the true coefficient vector β, unlike AMP or leave-one-out methods.
  • The estimate remains valid under general covariance matrices Σ ≠ I, removing the isotropy assumption common in prior high-dimensional analysis.
  • For the Huber Lasso, the estimate takes a closed form involving the active set size |Ŝ| and the norm of the projected residual, with explicit bounds on the error under sparsity and tuning parameter conditions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.