[Paper Review] When Does More Regularization Imply Fewer Degrees of Freedom? Sufficient Conditions and Counter Examples from Lasso and Ridge Regression
This paper challenges the common assumption that increased regularization reduces degrees of freedom in regression models. Using theoretical analysis and counterexamples, it demonstrates that in lasso and ridge regression, higher regularization can paradoxically increase degrees of freedom and optimism, undermining the core rationale for regularization. However, it identifies two cases—symmetric linear smoothers and convex-constrained linear regression—where monotonic reduction in degrees of freedom is guaranteed.
Regularization aims to improve prediction performance of a given statistical modeling approach by moving to a second approach which achieves worse training error but is expected to have fewer degrees of freedom, i.e., better agreement between training and prediction error. We show here, however, that this expected behavior does not hold in general. In fact, counter examples are given that show regularization can increase the degrees of freedom in simple situations, including lasso and ridge regression, which are the most common regularization approaches in use. In such situations, the regularization increases both training error and degrees of freedom, and is thus inherently without merit. On the other hand, two important regularization scenarios are described where the expected reduction in degrees of freedom is indeed guaranteed: (a) all symmetric linear smoothers, and (b) linear regression versus convex constrained linear regression (as in the constrained variant of ridge regression and lasso).
Motivation & Objective
- To investigate whether increased regularization consistently reduces degrees of freedom in regularized regression models.
- To challenge the widely held belief that higher regularization leads to lower optimism and better generalization.
- To identify sufficient conditions under which regularization does reduce degrees of freedom, and to expose cases where it does not.
- To provide counterexamples in lasso and ridge regression where more regularization increases both training error and degrees of freedom.
- To clarify the theoretical conditions under which the optimism of regularized models decreases monotonically with regularization.
Proposed method
- Formalizes nesting in regularization via projection mappings and defines monotonicity of degrees of freedom in terms of optimism and effective degrees of freedom.
- Uses Stein's lemma to express degrees of freedom as the trace of the Jacobian of the prediction function with respect to the observed data.
- Analyzes the Jacobian components of projection mappings onto nested subspaces to compare sensitivity of predictions to input perturbations.
- Applies non-expansion properties of projections onto convex sets to bound the influence of data perturbations on model outputs.
- Derives sufficient conditions under which the trace of the Jacobian (i.e., degrees of freedom) decreases with increased regularization.
- Constructs explicit counterexamples in lasso and ridge regression where higher regularization increases degrees of freedom and optimism.
Experimental results
Research questions
- RQ1Does increased regularization in lasso and ridge regression always lead to a reduction in degrees of freedom?
- RQ2Under what conditions can regularization increase optimism and degrees of freedom, contrary to theoretical expectations?
- RQ3Are there general classes of models where monotonic reduction in degrees of freedom with regularization is guaranteed?
- RQ4Why does the conventional intuition about regularization—reducing overfitting by lowering degrees of freedom—fail in lasso and ridge regression?
- RQ5Can the optimism of a regularized model be higher than that of a less regularized model, and if so, under what conditions?
Key findings
- In lasso and ridge regression, counterexamples exist where increased regularization leads to higher degrees of freedom and higher optimism, contradicting the standard assumption.
- For symmetric linear smoothers, increased regularization guarantees a monotonic decrease in degrees of freedom.
- In convex-constrained linear regression (e.g., constrained ridge and lasso), regularization also ensures a monotonic reduction in degrees of freedom.
- The optimism of a model is not necessarily monotonic with respect to the regularization parameter; higher regularization can increase optimism.
- The trace of the Jacobian matrix (effective degrees of freedom) can increase with regularization, even when training error increases, rendering such regularization ineffective.
- Theoretical conditions are derived under which the degrees of freedom of nested regularized models decrease monotonically with increased regularization.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.