[Paper Review] Model-free controlled variable selection via data splitting
This paper proposes a model-free variable selection method in sufficient dimension reduction using data splitting to control the false discovery rate (FDR) without assuming a specific regression model. By transforming responses and leveraging symmetry in test statistics derived from split data, the method achieves finite-sample and asymptotic FDR control while maintaining high power, outperforming existing approaches in simulations.
Addressing the simultaneous identification of contributory variables while controlling the false discovery rate (FDR) in high-dimensional data is a crucial statistical challenge. In this paper, we propose a novel model-free variable selection procedure in sufficient dimension reduction framework via a data splitting technique. The variable selection problem is first converted to a least squares procedure with several response transformations. We construct a series of statistics with global symmetry property and leverage the symmetry to derive a data-driven threshold aimed at error rate control. Our approach demonstrates the capability for achieving finite-sample and asymptotic FDR control under mild theoretical conditions. Numerical experiments confirm that our procedure has satisfactory FDR control and higher power compared with existing methods.
Motivation & Objective
- To develop a model-free variable selection procedure that controls the false discovery rate (FDR) in high-dimensional sufficient dimension reduction.
- To address the lack of error rate control in existing model-free SDR methods, particularly in finite samples.
- To enable effective identification of truly contributing variables while filtering out non-informative ones without parametric model assumptions.
- To achieve both finite-sample and asymptotic FDR control under mild regularity conditions.
- To improve power and accuracy compared to existing methods through a data-driven threshold derived from symmetry properties.
Proposed method
- The method transforms the variable selection problem into inference on regression coefficients via multiple response transformations in a least squares framework.
- It uses data splitting to create two independent samples, enabling the construction of symmetric test statistics for each covariate.
- Symmetric statistics are derived from the difference in estimated coefficients across the two data splits, ensuring global symmetry under the null hypothesis.
- A data-driven threshold is constructed using the symmetry property to control FDR without relying on p-values or parametric assumptions.
- The procedure is agnostic to the underlying SDR method, allowing integration with various existing SDR techniques like SIR or sliced inverse regression.
- FDR control is achieved via a knockoff-like thresholding rule based on the symmetry of test statistics, with theoretical justification under mild conditions.
Experimental results
Research questions
- RQ1Can a model-free variable selection procedure control the false discovery rate in sufficient dimension reduction without assuming a specific parametric model?
- RQ2How can symmetry in test statistics derived from data splitting be exploited to construct a data-driven threshold for FDR control?
- RQ3Does the proposed method achieve finite-sample and asymptotic FDR control under mild regularity conditions?
- RQ4How does the method's power compare to existing model-free and model-based variable selection procedures in high-dimensional settings?
- RQ5Can the method be flexibly combined with different existing SDR methods without compromising FDR control?
Key findings
- The proposed method achieves finite-sample and asymptotic FDR control under mild regularity conditions, including boundedness of design moments and sufficient signal strength.
- Theoretical analysis shows that the FDP (false discovery proportion) converges in probability to a value bounded by the nominal level α, ensuring FDR control.
- The method maintains high statistical power, as demonstrated by simulations showing better detection of true signals compared to existing approaches.
- The data-driven threshold derived from symmetry properties outperforms fixed thresholds and avoids reliance on p-values or asymptotic approximations.
- The method is robust to model misspecification and can be seamlessly integrated with various SDR techniques, enhancing interpretability by selecting only relevant predictors.
- Numerical experiments confirm that the procedure controls FDR at the nominal level while achieving higher power than competing methods in both low- and high-dimensional settings.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.