[Paper Review] Signal Adaptive Variable Selector for the Horseshoe Prior
This paper proposes the Signal Adaptive Variable Selector (SAVS), a fully automated, tuning-free method for variable selection in high-dimensional regression using the horseshoe prior. SAVS post-processes the posterior mean to identify signals and nulls by adaptively thresholding coefficients based on signal strength, achieving robust performance across diverse designs and outperforming frequentist and Bayesian alternatives in simulation and real genomic data with 20,000+ covariates.
In this article, we propose a simple method to perform variable selection as a post model-fitting exercise using continuous shrinkage priors such as the popular horseshoe prior. The proposed Signal Adaptive Variable Selector (SAVS) approach post-processes a point estimate such as the posterior mean to group the variables into signals and nulls. The approach is completely automated and does not require specification of any tuning parameters. We carried out a comprehensive simulation study to compare the performance of the proposed SAVS approach to frequentist penalization procedures and Bayesian model selection procedures. SAVS was found to be highly competitive across all the settings considered, and was particularly found to be robust to correlated designs. We also applied SAVS to a genomic dataset with more than 20,000 covariates to illustrate its scalability.
Motivation & Objective
- To address the lack of automated, tuning-parameter-free variable selection methods for continuous shrinkage priors like the horseshoe prior.
- To develop a post-processing approach that transforms posterior means into sparse estimates without requiring user-specified thresholds.
- To improve robustness to correlated designs compared to existing frequentist and Bayesian variable selection techniques.
- To demonstrate scalability of the method on high-dimensional genomic data with over 20,000 covariates.
Proposed method
- SAVS post-processes a point estimate such as the posterior mean or median obtained via MCMC or deterministic approximation.
- It adaptively thresholds the posterior mean using a signal-strength-based criterion that groups variables into signals and nulls.
- The method is fully automated and does not require tuning parameters, relying on the empirical distribution of posterior means to define signal boundaries.
- It leverages the posterior mean's shrinkage properties to distinguish between non-null (signal) and null (noise) coefficients.
- The approach is generalizable to other global-local shrinkage priors beyond the horseshoe, though the paper focuses on the horseshoe prior.
- The procedure is computationally efficient and scalable, enabling application to large-scale datasets such as those with p > 20,000 covariates.
Experimental results
Research questions
- RQ1Can a fully automated, tuning-free method be developed to perform variable selection after fitting a horseshoe prior model?
- RQ2How does SAVS perform in comparison to frequentist penalization methods like SCAD, MCP, and adaptive lasso in high-dimensional, correlated designs?
- RQ3Does SAVS maintain high power and low false discovery rate across varying signal strengths and correlation structures?
- RQ4Can SAVS be effectively scaled to large-scale genomic datasets with tens of thousands of covariates?
Key findings
- SAVS achieved a median true positive rate of 0.96 across all simulation settings, with 100% power in the orthogonal design (Set-1) and high power (0.81–0.94) in correlated designs (Set-2).
- In correlated designs (Set-2), SAVS maintained high true positive rates (0.80–0.95) and low false discovery rates (0.00–0.19), outperforming SCAD, adaptive lasso, and MCP in several configurations.
- For the genomic dataset with over 20,000 covariates, SAVS successfully identified relevant signals with high precision and scalability, demonstrating practical utility in real-world applications.
- SAVS was particularly robust to correlated designs, where frequentist methods like SCAD and MCP showed significant performance drops (e.g., 0.00–0.06 true positive rate in Set-2).
- The method achieved 100% true positive rate in 8 out of 10 simulation scenarios under Set-1 (orthogonal design), indicating strong performance in ideal conditions.
- In scenarios with weak signals, SAVS maintained a true positive rate of 0.81 in Set-2, while adaptive lasso and MCP dropped to 0.59 and 0.60, respectively, showing superior sensitivity to weak signals.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.