Skip to main content
QUICK REVIEW

[Paper Review] Understanding overfitting peaks in generalization error: Analytical risk curves for $l_2$ and $l_1$ penalized interpolation

Partha P. Mitra|arXiv (Cornell University)|Jun 9, 2019
Sparse and Compressive Sensing TechniquesEngineering24 references35 citations
TL;DR

The paper introduces MiSpaR (Misparametrized Sparse Regression) to analytically derive training and generalization error curves for $l_2$ and $l_1$ penalized interpolation in high-dimensional settings, highlighting that overfitting peaks do not strictly demarcate classical vs modern regimes and showing when each penalty generalizes well.

ABSTRACT

Traditionally in regression one minimizes the number of fitting parameters or uses smoothing/regularization to trade training (TE) and generalization error (GE). Driving TE to zero by increasing fitting degrees of freedom (dof) is expected to increase GE. However modern big-data approaches, including deep nets, seem to over-parametrize and send TE to zero (data interpolation) without impacting GE. Overparametrization has the benefit that global minima of the empirical loss function proliferate and become easier to find. These phenomena have drawn theoretical attention. Regression and classification algorithms have been shown that interpolate data but also generalize optimally. An interesting related phenomenon has been noted: the existence of non-monotonic risk curves, with a peak in GE with increasing dof. It was suggested that this peak separates a classical regime from a modern regime where over-parametrization improves performance. Similar over-fitting peaks were reported previously (statistical physics approach to learning) and attributed to increased fitting model flexibility. We introduce a generative and fitting model pair ("Misparametrized Sparse Regression" or MiSpaR) and show that the overfitting peak can be dissociated from the point at which the fitting function gains enough dof's to match the data generative model and thus provides good generalization. This complicates the interpretation of overfitting peaks as separating a "classical" from a "modern" regime. Data interpolation itself cannot guarantee good generalization: we need to study the interpolation with different penalty terms. We present analytical formulae for GE curves for MiSpaR with $l_2$ and $l_1$ penalties, in the interpolating limit $λ ightarrow 0$.These risk curves exhibit important differences and help elucidate the underlying phenomena.

Motivation & Objective

  • Introduce the Misparametrized Sparse Regression (MiSpaR) framework to separate measurements, model parameters, and fitting degrees of freedom.
  • Derive analytical expressions for training and generalization errors under $l_2$ and $l_1$ penalties in the interpolating limit.
  • Show how overfitting peaks relate to interpolation vs true data-generating capabilities and how sparsity and noise influence generalization.
  • Compare ridge ($l_2$) and sparse ($l_1$) penalties to illustrate when regularization improves or degrades generalization in overparameterized regimes.

Proposed method

  • Propose MiSpaR with a generative model where the number of inference parameters $p$ can differ from the generative parameter count $n$ and from measurements $m$.
  • Derive high-dimensional asymptotics for $m,p,n\to\infty$ with fixed ratios $\mu=p/m$ and $\alpha=m/n$ to obtain analytical TE and GE for $l_2$ regression.
  • Provide analytical GE expressions for $l_1$ penalties and present a pair of nonlinear equations for numerical solution in the interpolating limit.
  • Exhibit how effective noise is altered by undersampling/oversampling ($\alpha$, $\mu$) and sparsity ($\rho$) under both penalties.
  • Use self-averaging arguments and random matrix theory (Marchenko-Pastur distribution) to compute needed sums in GE/TE expressions.

Experimental results

Research questions

  • RQ1How do misparametrization and sparsity affect the training and generalization error when interpolating data with $l_2$ vs $l_1$ penalties?
  • RQ2Where do overfitting peaks occur in relation to the data interpolation point ($\mu=1$) and the regime of good generalization (e.g., $\mu\alpha=1$)?
  • RQ3How do the $l_2$ and $l_1$ penalties differ in their ability to generalize in highly overparameterized settings, especially under low noise and sparsity?
  • RQ4What are the precise analytical forms of GE and TE in the interpolating limit for both penalties, and how do these depend on $\alpha$, $\mu$, and $\rho$?

Key findings

  • In the interpolating limit ($\lambda\to 0$), the overfitting peak occurs at $\mu=1$ for both penalties, but good generalization can begin at $\mu\alpha=1$, not at the interpolation point.
  • For large overparameterization, generalization vanishes ($GE(\mu\to\infty)=1$) for both penalties, yet sparse $l_1$ can generalize well over a substantial $\mu$ range when $\sigma^2$ and $\rho$ are small.
  • There is a notable gap in performance between $l_1$ and $l_2$ in regimes of high overparameterization with small noise and strong sparsity, where $l_1$ can generalize while $l_2$ fails.
  • Finite regularization ($\lambda>0$) suppresses the overfitting peak, indicating interpolation alone does not guarantee good generalization.
  • Analytical GE expressions for $l_1$ involve a system of three equations linking $\tau$, $\hat{\rho}$, and $\sigma_{\xi}$, illustrating the algorithmic phase transitions in sparse regression.
  • The work demonstrates that generalization properties depend strongly on inductive bias (choice of penalty) and are not intrinsic to data interpolation alone.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.