[Paper Review] Asymptotic Properties of Lasso+mLS and Lasso+Ridge in Sparse High-dimensional Linear Regression
This paper establishes the asymptotic unbiasedness, normality, and oracle properties of Lasso+mLS and Lasso+Ridge estimators in high-dimensional sparse linear models. It proposes a parametric residual bootstrap procedure for valid inference and proves that both methods achieve the optimal $s/n$ convergence rate for mean squared error with bias decaying exponentially fast.
We study the asymptotic properties of Lasso+mLS and Lasso+Ridge under the sparse high-dimensional linear regression model: Lasso selecting predictors and then modified Least Squares (mLS) or Ridge estimating their coefficients. First, we propose a valid inference procedure for parameter estimation based on parametric residual bootstrap after Lasso+mLS and Lasso+Ridge. Second, we derive the asymptotic unbiasedness of Lasso+mLS and Lasso+Ridge. More specifically, we show that their biases decay at an exponential rate and they can achieve the oracle convergence rate of $s/n$ (where $s$ is the number of nonzero regression coefficients and $n$ is the sample size) for mean squared error (MSE). Third, we show that Lasso+mLS and Lasso+Ridge are asymptotically normal. They have an oracle property in the sense that they can select the true predictors with probability converging to 1 and the estimates of nonzero parameters have the same asymptotic normal distribution that they would have if the zero parameters were known in advance. In fact, our analysis is not limited to adopting Lasso in the selection stage, but is applicable to any other model selection criteria with exponentially decay rates of the probability of selecting wrong models.
Motivation & Objective
- To establish the asymptotic unbiasedness and normality of Lasso+mLS and Lasso+Ridge estimators in high-dimensional sparse linear models.
- To develop a valid inference procedure using parametric residual bootstrap after Lasso-based selection.
- To show that both methods achieve the oracle convergence rate of $s/n$ for mean squared error.
- To prove that the estimators possess the oracle property: consistent model selection and asymptotically normal estimates as if the true model were known.
- To extend the analysis beyond Lasso to any model selection criterion with exponentially decaying error rates in model selection.
Proposed method
- Proposes a parametric residual bootstrap procedure to approximate the sampling distribution of Lasso+mLS and Lasso+Ridge estimators.
- Analyzes the asymptotic behavior under the high-dimensional setting where $p$ and $s$ grow with $n$, with $s$ being the number of nonzero coefficients.
- Uses the Irrepresentable Condition and restricted eigenvalue-type assumptions to ensure model selection consistency.
- Applies concentration inequalities and metric-based bounds (e.g., $d^2$) to control the discrepancy between empirical and true distributions of residuals.
- Derives convergence rates for estimation error using $\ell_2$ and $\ell_1$ norms under weak regularity conditions.
- Establishes asymptotic normality by verifying that the conditions for the central limit theorem hold in the high-dimensional regime after model selection.
Experimental results
Research questions
- RQ1Can Lasso+mLS and Lasso+Ridge achieve asymptotic unbiasedness in high-dimensional sparse linear models?
- RQ2Do these two-stage estimators attain the oracle convergence rate of $s/n$ for mean squared error?
- RQ3Are Lasso+mLS and Lasso+Ridge asymptotically normal, with the same limiting distribution as if the true model were known?
- RQ4Can a residual bootstrap procedure provide valid inference for these two-stage estimators?
- RQ5Does the analysis extend to other model selection methods beyond Lasso, provided they have exponentially decaying error rates in selecting the wrong model?
Key findings
- Lasso+mLS and Lasso+Ridge are asymptotically unbiased, with bias decaying at an exponential rate as the sample size increases.
- Both methods achieve the optimal $s/n$ convergence rate for mean squared error, matching the oracle rate.
- The estimators are asymptotically normal, with the same limiting distribution as if the true support of $\beta^*$ were known in advance.
- The proposed parametric residual bootstrap procedure provides valid inference for the two-stage estimators under regularity conditions.
- The oracle property holds: the probability of selecting the correct model converges to one, and the estimates of nonzero coefficients are asymptotically normal.
- The theoretical framework applies not only to Lasso but to any model selection criterion with exponentially decaying model selection error rates.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.