[Paper Review] High-dimensional simultaneous inference with the bootstrap
This paper proposes a residual and wild bootstrap methodology for high-dimensional linear models with non-Gaussian and heteroscedastic errors, using the de-sparsified Lasso as a foundation. It establishes asymptotic consistency for simultaneous inference under sparsity and group size constraints, offering robust, computationally feasible confidence intervals and multiple testing adjustments in high-dimensional settings where $ p \gg n $.
We propose a residual and wild bootstrap methodology for individual and simultaneous inference in high-dimensional linear models with possibly non-Gaussian and heteroscedastic errors. We establish asymptotic consistency for simultaneous inference for parameters in groups $G$, where $p \gg n$, $s_0 = o(n^{1/2}/\{\log(p) \log(|G|)^{1/2}\})$ and $\log(|G|) = o(n^{1/7})$, with $p$ the number of variables, $n$ the sample size and $s_0$ denoting the sparsity. The theory is complemented by many empirical results. Our proposed procedures are implemented in the R-package hdi.
Motivation & Objective
- Address the lack of reliable inference methods in high-dimensional linear models with non-Gaussian and heteroscedastic errors.
- Develop a bootstrap-based approach that avoids the super-efficiency bias inherent in standard Lasso-based inference.
- Enable simultaneous inference for groups of variables under high-dimensional asymptotics with $ p \gg n $.
- Establish theoretical consistency for both individual and simultaneous inference under weak sparsity and group size constraints.
- Provide a practical, computationally efficient method implemented in the R package hdi for real-world high-dimensional data analysis.
Proposed method
- The paper proposes bootstrapping the de-sparsified Lasso estimator, which is a regular, non-sparse estimator with asymptotic normality and efficiency under weak sparsity.
- Residual and wild bootstrap schemes are employed to approximate the sampling distribution of the de-sparsified Lasso, with the wild bootstrap preferred for heteroscedastic errors.
- Bootstrap resampling is applied to the residual vectors $ Z_1, \dots, Z_p $, which are computed once for fixed design, minimizing computational overhead.
- Theoretical consistency is established under the conditions $ s_0 = o(n^{1/2}/\{\log(p)\log(|G|)^{1/2}\}) $ and $ \log(|G|) = o(n^{1/7}) $, where $ s_0 $ is sparsity and $ G $ is a group of variables.
- Multiple testing adjustment is achieved through bootstrap-based p-values, enabling simultaneous error control across many hypotheses.
- Implementation is integrated into the R package hdi (Meier et al., 2016), supporting practical use in high-dimensional data analysis.
Experimental results
Research questions
- RQ1How can reliable simultaneous inference be achieved in high-dimensional linear models with $ p \gg n $ and non-Gaussian, heteroscedastic errors?
- RQ2What bootstrap method—residual or wild—provides consistent inference under heteroscedasticity, and why is one preferred over the other?
- RQ3Can the bootstrap of the de-sparsified Lasso avoid the super-efficiency bias that plagues direct Lasso-based inference?
- RQ4How do the theoretical conditions on sparsity $ s_0 $ and group size $ |G| $ affect the validity of simultaneous inference?
- RQ5What is the computational feasibility and empirical performance of the proposed bootstrap method compared to existing alternatives?
Key findings
- The proposed bootstrap methodology achieves asymptotic consistency for simultaneous inference under the conditions $ s_0 = o(n^{1/2}/\{\log(p)\log(|G|)^{1/2}\}) $ and $ \log(|G|) = o(n^{1/7}) $, even when $ p \gg n $.
- The wild bootstrap is consistent under heteroscedastic errors, while the residual bootstrap fails in this setting, making the wild bootstrap the preferred method for robust inference.
- By bootstrapping the de-sparsified Lasso instead of the sparse Lasso, the method avoids the super-efficiency phenomenon, leading to more reliable confidence intervals and p-values.
- Empirical results demonstrate superior performance in multiple testing adjustment and simultaneous coverage compared to alternative methods, especially under non-Gaussian and heteroscedastic errors.
- The computational cost of bootstrapping is manageable and not substantially higher than computing the de-sparsified Lasso itself, especially for fixed design.
- The procedures are implemented in the R package hdi (Meier et al., 2016), enabling practical application to real-world high-dimensional datasets.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.