[Paper Review] Degrees of Freedom and Model Search
This paper derives an exact expression for the degrees of freedom of best subset selection under orthogonal predictors, introducing the concept of 'search degrees of freedom' to quantify the effective cost of model selection. It extends Stein's formula to handle discontinuous functions, showing that for best subset selection, degrees of freedom exceed the number of selected variables due to adaptive selection, unlike the lasso where degrees of freedom equal the expected number of selected variables.
Degrees of freedom is a fundamental concept in statistical modeling, as it provides a quantitative description of the amount of fitting performed by a given procedure. But, despite this fundamental role in statistics, its behavior not completely well-understood, even in some fairly basic settings. For example, it may seem intuitively obvious that the best subset selection fit with subset size k has degrees of freedom larger than k, but this has not been formally verified, nor has is been precisely studied. In large part, the current paper is motivated by this particular problem, and we derive an exact expression for the degrees of freedom of best subset selection in a restricted setting (orthogonal predictor variables). Along the way, we develop a concept that we name "search degrees of freedom"; intuitively, for adaptive regression procedures that perform variable selection, this is a part of the (total) degrees of freedom that we attribute entirely to the model selection mechanism. Finally, we establish a modest extension of Stein's formula to cover discontinuous functions, and discuss its potential role in degrees of freedom and search degrees of freedom calculations.
Motivation & Objective
- To formally characterize the degrees of freedom of best subset selection, a procedure where intuitive expectations about overfitting are not yet rigorously established.
- To introduce and formalize the concept of 'search degrees of freedom'—the portion of total degrees of freedom attributable solely to the model selection mechanism in adaptive regression.
- To extend Stein's formula to cover discontinuous functions, enabling degrees of freedom calculations in non-smooth, adaptive procedures like best subset selection.
- To contrast the behavior of degrees of freedom in best subset selection with that of the lasso, where degrees of freedom equal the expected number of selected variables due to shrinkage.
Proposed method
- Derives an exact expression for degrees of freedom of best subset selection under orthogonal predictor variables using a novel extension of Stein's formula to discontinuous functions.
- Introduces the concept of 'search degrees of freedom' as the component of total degrees of freedom arising purely from the model selection process.
- Uses integration by parts and conditional expectation techniques to handle the non-smoothness of the selection mechanism in best subset selection.
- Applies the extended Stein's formula to compute the expected sensitivity of the fitted values to changes in the response, which directly yields the degrees of freedom.
- Establishes that for best subset selection with orthogonal predictors, degrees of freedom exceed the number of selected variables due to the adaptive search over model space.
- Compares the derived result with the known lasso degrees of freedom formula, highlighting the role of shrinkage in balancing selection cost.
Experimental results
Research questions
- RQ1Does best subset selection with subset size $k$ have degrees of freedom strictly greater than $k$, and if so, by how much?
- RQ2What is the precise contribution of the model selection mechanism to the total degrees of freedom in adaptive regression procedures?
- RQ3How can Stein's formula be extended to handle discontinuous functions arising in model selection procedures like best subset selection?
- RQ4Why does the lasso have degrees of freedom equal to the expected number of selected variables, while best subset selection does not?
- RQ5Can the concept of 'search degrees of freedom' be formally defined and quantified in a way that isolates the cost of variable selection from estimation?
Key findings
- For best subset selection with orthogonal predictors, the degrees of freedom is strictly greater than $k$, the number of selected variables, due to the adaptive selection process.
- The paper derives an exact analytical expression for the degrees of freedom of best subset selection under orthogonal design, confirming that it exceeds $k$.
- The concept of 'search degrees of freedom' is formally introduced and quantified, representing the effective cost of model search independent of estimation.
- The extended Stein's formula allows for degrees of freedom calculations in non-smooth, discontinuous procedures like best subset selection.
- Unlike best subset selection, the lasso's degrees of freedom equal the expected number of selected variables, due to the balancing effect of $ abla_1$ shrinkage.
- The result shows that the cost of model search in best subset selection is non-trivial and must be explicitly accounted for in model comparison and risk estimation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.