Skip to main content
QUICK REVIEW

[Paper Review] Extended BIC for linear regression models with diverging number of relevant features and high or ultra-high feature spaces

Shan Luo, Zehua Chen|arXiv (Cornell University)|Jul 13, 2011
Statistical Methods and Inference19 references3 citations
TL;DR

This paper extends the Bayes Information Criterion (BIC) to linear regression models with high or ultra-high-dimensional feature spaces where the number of relevant features diverges with sample size. It establishes selection consistency of the Extended BIC (EBIC) under general conditions, proving that EBIC consistently identifies the true model when the number of relevant features grows at rate $ O(n^c) $ for $ 0 < c < 1 $, even as $ p = O(n^k) $ or $ p = O(e^{n^k}) $.

ABSTRACT

In many conventional scientific investigations with high or ultra-high dimensional feature spaces, the relevant features, though sparse, are large in number compared with classical statistical problems, and the magnitude of their effects tapers off. It is reasonable to model the number of relevant features as a diverging sequence when sample size increases. In this article, we investigate the properties of the extended Bayes information criterion (EBIC) (Chen and Chen, 2008) for feature selection in linear regression models with diverging number of relevant features in high or ultra-high dimensional feature spaces. The selection consistency of the EBIC in this situation is established. The application of EBIC to feature selection is considered in a two-stage feature selection procedure. Simulation studies are conducted to demonstrate the performance of the EBIC together with the two-stage feature selection procedure in finite sample cases.

Motivation & Objective

  • To address the lack of model selection criteria that maintain selection consistency in high or ultra-high-dimensional linear models with a diverging number of relevant features.
  • To establish theoretical conditions under which the Extended BIC (EBIC) remains selection consistent when the number of true features increases with sample size.
  • To extend the applicability of EBIC beyond fixed or slowly growing numbers of relevant features to settings where the number of relevant features grows polynomially or exponentially with sample size.
  • To provide a theoretically grounded criterion for selecting the penalty parameter in penalized likelihood methods in small-$ n $-large-$ p $ problems.
  • To validate the finite-sample performance of EBIC through simulation studies in two-stage feature selection procedures.

Proposed method

  • Proposes an extended version of the BIC (EBIC) indexed by a parameter $ \gamma \in [0,1] $, which penalizes model complexity based on the number of features and the logarithm of the number of possible models.
  • Derives selection consistency of EBIC under the assumption that the number of relevant features $ p_{0n} = O(n^c) $ for $ 0 < c < 1 $, with $ p_n $ growing polynomially or exponentially in $ n $.
  • Uses asymptotic analysis of the likelihood ratio and residual sum of squares to compare the EBIC values of the true model $ s_{0n} $ and any alternative model $ s $, showing that EBIC favors the true model with probability approaching one.
  • Applies concentration inequalities and chi-squared tail bounds to control the stochastic behavior of the residual sum of squares under the null and alternative models.
  • Introduces a two-stage feature selection procedure: first, a screening step reduces the feature space; second, EBIC is used to select the optimal model from the reduced set.
  • Employs theoretical bounds on the maximum eigenvalues and quadratic forms of projection matrices to control the deviation of the residual sum of squares from its expectation.

Experimental results

Research questions

  • RQ1Under what conditions is the Extended BIC (EBIC) selection consistent when the number of relevant features diverges with sample size in high or ultra-high-dimensional linear models?
  • RQ2How does the EBIC parameter $ \gamma $ affect selection consistency when the number of true features grows as $ O(n^c) $ for $ 0 < c < 1 $?
  • RQ3Can EBIC maintain selection consistency when the number of features $ p_n $ grows at polynomial or exponential rates relative to sample size $ n $?
  • RQ4How does EBIC compare to traditional criteria like AIC, BIC, and CV in terms of selection consistency in small-$ n $-large-$ p $ problems with diverging true feature counts?
  • RQ5What is the finite-sample performance of EBIC when used in a two-stage feature selection procedure?

Key findings

  • The EBIC is selection consistent when $ p_n = O(n^k) $ or $ p_n = O(e^{n^k}) $ for $ 0 < k < 1 $, and the number of relevant features grows as $ O(n^c) $ with $ 0 < c < 1 $, provided $ \gamma > \frac{1+\delta}{1-\delta} - \frac{\ln n}{2(1-\delta)\ln p_n} $, where $ \delta = \lim \frac{\ln \nu(s)}{\ln p_n} $.
  • The EBIC outperforms traditional criteria like AIC, BIC, and CV in terms of selection consistency in high-dimensional settings with diverging numbers of relevant features.
  • The EBIC maintains selection consistency even when the number of relevant features increases with sample size, which is not guaranteed by classical BIC or AIC.
  • Simulation studies confirm that the EBIC-based two-stage feature selection procedure achieves high true positive and low false positive rates in finite samples.
  • The theoretical analysis shows that the difference between EBIC values of the true model and any alternative model diverges to infinity in probability, ensuring consistent model selection.
  • The EBIC with $ \gamma = 1 $ asymptotically recovers the mBIC criterion used in genetic QTL mapping, validating its consistency in a well-known application context.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.