[Paper Review] Regularization after retention in ultrahigh dimensional linear regression models
This paper proposes a novel two-step variable selection framework for ultrahigh-dimensional linear regression, where significant variables are retained based on marginal correlations before applying regularization only to the remaining variables. The method achieves model selection consistency and outperforms independence screening and Lasso in scenarios with weak marginal signals, especially in finite samples.
In ultrahigh dimensional setting, independence screening has been both theoretically and empirically proved a useful variable selection framework with low computation cost. In this work, we propose a two-step framework by using marginal information in a different perspective from independence screening. In particular, we retain significant variables rather than screening out irrelevant ones. The new method is shown to be model selection consistent in the ultrahigh dimensional linear regression model. To improve the finite sample performance, we then introduce a three-step version and characterize its asymptotic behavior. Simulations and real data analysis show advantages of our method over independence screening and its iterative variants in certain regimes.
Motivation & Objective
- To address the limitation of independence screening in ultrahigh-dimensional settings where important variables have weak marginal correlations.
- To improve finite-sample performance in variable selection when signals are weak or correlated with noise.
- To develop a theoretically grounded method that ensures model selection consistency under weaker assumptions than traditional screening methods.
- To provide a principled alternative to screening-out approaches by focusing on retention of potentially important variables.
- To demonstrate the superiority of the proposed method over Lasso and iterative screening variants in specific simulation regimes.
Proposed method
- The method uses marginal regression coefficient estimates to retain a set of predictors in the first step, based on a permutation-based thresholding procedure.
- In the second step, regularization via penalized least squares is applied only to the variables not retained, reducing the dimensionality of the optimization problem.
- Theoretical analysis replaces the standard lower-bound assumption on true signals with an upper-bound assumption on irrelevant variables’ marginal correlations.
- A three-step extension is introduced to correct for false positives in the retention step, enhancing finite-sample performance.
- The method is shown to be model selection consistent under mild regularity conditions, with asymptotic properties derived for both two- and three-step versions.
- Theoretical comparison with Lasso shows improved selection consistency under certain conditions, particularly when signals are weak.
Experimental results
Research questions
- RQ1Can a retention-based approach outperform traditional screening methods in ultrahigh-dimensional linear models with weak marginal signals?
- RQ2Does the proposed two-step framework achieve model selection consistency when independence screening fails due to weak marginal correlations?
- RQ3How does the three-step extension improve finite-sample performance by correcting false positives from the retention step?
- RQ4What are the theoretical conditions under which the proposed method achieves selection consistency, and how do they compare to those of Lasso or SIS?
- RQ5In what regimes does the proposed method demonstrate superior performance compared to Lasso and iterative screening variants?
Key findings
- The proposed method achieves model selection consistency in ultrahigh-dimensional linear regression, even when important variables have weak marginal correlations.
- In simulations, the method (RAR+) significantly outperforms SIS-lasso and ISIS-lasso in terms of selection accuracy, especially in scenarios with weak signals.
- For Scenario 4(C), the RAR+ method achieved a mean signal recovery proportion of 0.81 (SD 0.08) at (n,p)=(700,4073), compared to 0.56 for SIS-lasso.
- The three-step version (RAR+) reduced model size from over 100 to around 20 in high-dimensional settings, indicating effective elimination of false positives.
- In Scenario 4(D), the RAR+(MC+) method achieved a model size of 20.50 (SD 0.95) at (n,p)=(700,4073), closely matching the true model size of 20.
- Theoretical analysis shows that the method’s consistency condition is weaker than that of Lasso, particularly when signals are weak or correlated with noise.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.