Skip to main content
QUICK REVIEW

[Paper Review] Selective inference with unknown variance via the square-root LASSO

Xiaoying Tian, Joshua R. Loftus|arXiv (Cornell University)|Apr 29, 2015
Statistical Methods and Inference12 references16 citations
TL;DR

This paper proposes a selective inference framework using the square-root LASSO to simultaneously estimate regression coefficients, their confidence intervals, and the error variance $σ^2$ after model selection, even when $σ^2$ is unknown. The method provides valid p-values and confidence intervals by leveraging the scale-free tuning property of the square-root LASSO, which eliminates dependence on unknown noise levels, and achieves superior variance estimation and robust FDR control in simulations compared to existing methods.

ABSTRACT

There has been much recent work on inference after model selection when the noise level is known, however, $σ$ is rarely known in practice and its estimation is difficult in high-dimensional settings. In this work we propose using the square-root LASSO (also known as the scaled LASSO) to perform selective inference for the coefficients and the noise level simultaneously. The square-root LASSO has the property that choosing a reasonable tuning parameter is scale-free, namely it does not depend on the noise level in the data. We provide valid p-values and confidence intervals for the coefficients after selection, and estimates for model specific variance. Our estimates perform better than other estimates of $σ^2$ in simulation.

Motivation & Objective

  • To address the challenge of selective inference in high-dimensional regression when the error variance $σ^2$ is unknown and difficult to estimate.
  • To develop a method that jointly performs inference on regression coefficients and estimates the error variance $σ^2$ after model selection.
  • To ensure valid statistical inference (p-values, confidence intervals) that accounts for data-dependent model selection, even when $σ^2$ is unknown.
  • To improve estimation accuracy of $σ^2$ compared to existing methods like those in Sun & Zhang (2011) and Reid et al. (2013).
  • To enable robust false discovery rate (FDR) control in multiple testing scenarios under model selection.

Proposed method

  • Uses the square-root LASSO as the model selection procedure, which has a tuning parameter $λ$ that is invariant to the noise level $σ^2$ due to its scale-free property.
  • Derives the Karush-Kuhn-Tucker (KKT) conditions for the square-root LASSO to characterize the selection event and define the selective inference framework.
  • Applies the selective inference framework of Fithian et al. (2014) to construct valid p-values and confidence intervals conditional on the selected model and signs.
  • Approximates the distribution of test statistics under the selective distribution using a truncated Gaussian or t-distribution, with variance estimated via $σ_P^2(y)$.
  • Uses the scaled residual vector $U_{-E}(y)$, which is ancillary under the null, to perform regression diagnostics and goodness-of-fit tests.
  • Applies the Benjamini-Hochberg procedure to the derived p-values for FDR control in multiple hypothesis testing.

Experimental results

Research questions

  • RQ1Can selective inference be performed reliably when the error variance $σ^2$ is unknown and must be estimated from the same data used for model selection?
  • RQ2Does the square-root LASSO enable valid inference on coefficients and $σ^2$ simultaneously, with tuning independent of $σ^2$?
  • RQ3How does the proposed method compare to existing variance estimators in terms of accuracy and robustness in high-dimensional settings?
  • RQ4Can the derived p-values support effective FDR control in multiple testing under model selection?
  • RQ5How robust is the method to variations in the tuning parameter $λ$ and design correlation $ρ$?

Key findings

  • The proposed method produces valid p-values and confidence intervals for regression coefficients after model selection, even when $σ^2$ is unknown.
  • The square-root LASSO’s tuning parameter $λ$ is independent of $σ^2$, enabling stable and scale-free model selection.
  • The method provides more accurate estimates of $σ^2$ than existing methods such as those in Sun & Zhang (2011) and Reid et al. (2013), as shown in simulations.
  • The FDR control is robust across different correlation levels $ρ$, with the method maintaining FDR at approximately 0.2 across all $ρ$ values tested.
  • The method achieves approximately 20% higher statistical power compared to the procedure in Barber & Candes (2016), particularly under high correlation designs.
  • Coverage of confidence intervals based on the Gaussian approximation is close to nominal levels (e.g., 95% interval covers 94.7% of the time), validating the approximation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.