Skip to main content
QUICK REVIEW

[Paper Review] Model-free controlled variable selection via data splitting

Yixin Han, Xu Guo|arXiv (Cornell University)|Oct 22, 2022
Control Systems and Identification4 citations
TL;DR

This paper proposes a model-free variable selection method in sufficient dimension reduction using data splitting to control the false discovery rate (FDR) without assuming a specific regression model. By transforming responses and leveraging symmetry in test statistics derived from split data, the method achieves finite-sample and asymptotic FDR control while maintaining high power, outperforming existing approaches in simulations.

ABSTRACT

Addressing the simultaneous identification of contributory variables while controlling the false discovery rate (FDR) in high-dimensional data is a crucial statistical challenge. In this paper, we propose a novel model-free variable selection procedure in sufficient dimension reduction framework via a data splitting technique. The variable selection problem is first converted to a least squares procedure with several response transformations. We construct a series of statistics with global symmetry property and leverage the symmetry to derive a data-driven threshold aimed at error rate control. Our approach demonstrates the capability for achieving finite-sample and asymptotic FDR control under mild theoretical conditions. Numerical experiments confirm that our procedure has satisfactory FDR control and higher power compared with existing methods.

Motivation & Objective

  • To develop a model-free variable selection procedure that controls the false discovery rate (FDR) in high-dimensional sufficient dimension reduction.
  • To address the lack of error rate control in existing model-free SDR methods, particularly in finite samples.
  • To enable effective identification of truly contributing variables while filtering out non-informative ones without parametric model assumptions.
  • To achieve both finite-sample and asymptotic FDR control under mild regularity conditions.
  • To improve power and accuracy compared to existing methods through a data-driven threshold derived from symmetry properties.

Proposed method

  • The method transforms the variable selection problem into inference on regression coefficients via multiple response transformations in a least squares framework.
  • It uses data splitting to create two independent samples, enabling the construction of symmetric test statistics for each covariate.
  • Symmetric statistics are derived from the difference in estimated coefficients across the two data splits, ensuring global symmetry under the null hypothesis.
  • A data-driven threshold is constructed using the symmetry property to control FDR without relying on p-values or parametric assumptions.
  • The procedure is agnostic to the underlying SDR method, allowing integration with various existing SDR techniques like SIR or sliced inverse regression.
  • FDR control is achieved via a knockoff-like thresholding rule based on the symmetry of test statistics, with theoretical justification under mild conditions.

Experimental results

Research questions

  • RQ1Can a model-free variable selection procedure control the false discovery rate in sufficient dimension reduction without assuming a specific parametric model?
  • RQ2How can symmetry in test statistics derived from data splitting be exploited to construct a data-driven threshold for FDR control?
  • RQ3Does the proposed method achieve finite-sample and asymptotic FDR control under mild regularity conditions?
  • RQ4How does the method's power compare to existing model-free and model-based variable selection procedures in high-dimensional settings?
  • RQ5Can the method be flexibly combined with different existing SDR methods without compromising FDR control?

Key findings

  • The proposed method achieves finite-sample and asymptotic FDR control under mild regularity conditions, including boundedness of design moments and sufficient signal strength.
  • Theoretical analysis shows that the FDP (false discovery proportion) converges in probability to a value bounded by the nominal level α, ensuring FDR control.
  • The method maintains high statistical power, as demonstrated by simulations showing better detection of true signals compared to existing approaches.
  • The data-driven threshold derived from symmetry properties outperforms fixed thresholds and avoids reliance on p-values or asymptotic approximations.
  • The method is robust to model misspecification and can be seamlessly integrated with various SDR techniques, enhancing interpretability by selecting only relevant predictors.
  • Numerical experiments confirm that the procedure controls FDR at the nominal level while achieving higher power than competing methods in both low- and high-dimensional settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.