Skip to main content
QUICK REVIEW

[Paper Review] Which bridge estimator is optimal for variable selection?

Shuaiwen Wang, Haolei Weng|arXiv (Cornell University)|May 24, 2017
Statistical Methods and Inference56 references13 citations
TL;DR

This paper investigates two-stage variable selection (TVS) using bridge estimators in high-dimensional linear models, where the number of predictors $ p $ grows proportionally to sample size $ n $. It establishes that the asymptotic false discovery proportion (AFDP) is minimized when the first-stage bridge estimator has the smallest asymptotic mean square error (AMSE), revealing that Ridge outperforms LASSO in high-noise regimes and that LASSO is suboptimal for rare, weak signals.

ABSTRACT

We study the problem of variable selection for linear models under the high-dimensional asymptotic setting, where the number of observations $n$ grows at the same rate as the number of predictors $p$. We consider two-stage variable selection techniques (TVS) in which the first stage uses bridge estimators to obtain an estimate of the regression coefficients, and the second stage simply thresholds this estimate to select the "important" predictors. The asymptotic false discovery proportion (AFDP) and true positive proportion (ATPP) of these TVS are evaluated. We prove that for a fixed ATPP, in order to obtain a smaller AFDP, one should pick a bridge estimator with smaller asymptotic mean square error in the first stage of TVS. Based on such principled discovery, we present a sharp comparison of different TVS, via an in-depth investigation of the estimation properties of bridge estimators. Rather than "order-wise" error bounds with loose constants, our analysis focuses on precise error characterization. Various interesting signal-to-noise ratio and sparsity settings are studied. Our results offer new and thorough insights into high-dimensional variable selection. For instance, we prove that a TVS with Ridge in its first stage outperforms TVS with other bridge estimators in large noise settings; two-stage LASSO becomes inferior when the signal is rare and weak. As a by-product, we show that two-stage methods outperform some standard variable selection techniques, such as LASSO and Sure Independence Screening, under certain conditions.

Motivation & Objective

  • To determine which bridge estimator yields optimal variable selection performance in high-dimensional linear models.
  • To analyze the impact of signal-to-noise ratio (SNR) and sparsity on the optimal choice of bridge estimator.
  • To establish a principled link between first-stage estimation accuracy (AMSE) and second-stage variable selection performance (AFDP, ATPP).
  • To compare TVS with standard methods like LASSO and Sure Independence Screening under asymptotic high-dimensional regimes.

Proposed method

  • The study employs a two-stage variable selection (TVS) framework: first, a bridge estimator $ \hat{\beta}(q,\lambda) $ is computed via $ \ell_q $-regularized least squares; second, the estimate is hard-thresholded to select variables.
  • The asymptotic false discovery proportion (AFDP) and true positive proportion (ATPP) are derived under the high-dimensional asymptotic regime $ n/p \to \delta \in (0,\infty) $.
  • The analysis relies on precise asymptotic characterizations of the bridge estimator's mean square error (AMSE), avoiding loose order-wise bounds.
  • Asymptotic distributions of the estimator and thresholded output are derived using tools from random matrix theory and approximate message passing (AMP).
  • The optimal bridge estimator is identified as the one minimizing AMSE in the first stage, which in turn minimizes AFDP for a fixed ATPP.
  • The framework is applied to three key scenarios: rare signals, large noise, and large sample sizes, with analytical results derived for each.

Experimental results

Research questions

  • RQ1Which bridge estimator minimizes the asymptotic false discovery proportion (AFDP) in high-dimensional variable selection?
  • RQ2Does LASSO outperform two-stage methods based on other bridge estimators, particularly in low signal-to-noise or rare signal regimes?
  • RQ3How does the signal-to-noise ratio (SNR) influence the optimal choice of the bridge parameter $ q $?
  • RQ4Under what conditions do two-stage methods outperform standard techniques like LASSO and Sure Independence Screening?
  • RQ5What is the precise relationship between the estimation error (AMSE) of the first-stage bridge estimator and the variable selection performance (AFDP, ATPP)?

Key findings

  • Ridge regression (q=2) is optimal among all bridge estimators in high-noise settings, as it minimizes the asymptotic mean square error (AMSE).
  • For rare and weak signals, LASSO (q=1) performs poorly; other bridge estimators with q>1 can outperform it when signal strength is below a critical threshold.
  • Two-stage methods based on bridge estimators outperform standard techniques like LASSO and Sure Independence Screening under certain high-dimensional asymptotic conditions.
  • The optimal bridge estimator for variable selection is the one with the smallest AMSE in the first stage, establishing a direct link between estimation accuracy and selection performance.
  • In the large sample regime, the results align with classical low-dimensional asymptotic theory, validating the consistency of the high-dimensional framework.
  • The AFDP and ATPP of TVS are precisely characterized via asymptotic distributions of the thresholded estimator, enabling accurate comparison across different bridge estimators.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.