Skip to main content
QUICK REVIEW

[论文解读] Variable Selection for Highly Correlated Predictors

Fei Xue, Annie Qu|arXiv (Cornell University)|Sep 14, 2017
Statistical Methods and Inference参考文献 39被引用 8
一句话总结

本文提出了一种新型变量选择方法——半标准化部分相关系数(SPAC),该方法在保留系数大小的同时减少了预测变量之间的相关性影响,从而在高度多重共线性条件下实现一致的变量选择。SPAC-Lasso在低维和高维设置下均实现了符号一致性,且在模拟实验和HapMap基因数据集分析中优于Lasso、自适应Lasso、SCAD和PC-simple方法。

ABSTRACT

Penalty-based variable selection methods are powerful in selecting relevant covariates and estimating coefficients simultaneously. However, variable selection could fail to be consistent when covariates are highly correlated. The partial correlation approach has been adopted to solve the problem with correlated covariates. Nevertheless, the restrictive range of partial correlation is not effective for capturing signal strength for relevant covariates. In this paper, we propose a new Semi-standard PArtial Covariance (SPAC) which is able to reduce correlation effects from other predictors while incorporating the magnitude of coefficients. The proposed SPAC variable selection facilitates choosing covariates which have direct association with the response variable, via utilizing dependency among covariates. We show that the proposed method with the Lasso penalty (SPAC-Lasso) enjoys strong sign consistency in both finite-dimensional and high-dimensional settings under regularity conditions. Simulation studies and the `HapMap' gene data application show that the proposed method outperforms the traditional Lasso, adaptive Lasso, SCAD, and Peter-Clark-simple (PC-simple) methods for highly correlated predictors.

研究动机与目标

  • 解决当预测变量高度相关时,传统基于惩罚的方法在选择相关变量方面失效的问题。
  • 克服部分相关性方法的局限性,后者在相关设置下限制了信号强度的捕捉能力。
  • 开发一种利用预测变量间依赖关系的方法,以识别与响应变量具有直接关联的协变量。
  • 在有限维和高维设置下确保变量选择的符号一致性。
  • 在具有强预测变量相关性的现实基因组数据(如HapMap基因数据)中提升性能。

提出的方法

  • 提出半标准化部分相关系数(SPAC),一种可减少其他预测变量相关性影响并同时保留系数大小的变换方法。
  • 对SPAC变换后的预测变量应用Lasso惩罚,从而得到SPAC-Lasso,实现变量选择与系数估计的同步进行。
  • 利用预测变量间的依赖结构,以分离出与响应变量的直接关联。
  • 采用正则性条件,建立SPAC-Lasso在有限维和高维设置下的理论一致性。
  • 结合部分协方差概念,但将其扩展以保留信号强度,克服经典部分相关性范围受限的问题。
  • 引入标准化步骤,以确保尺度不变性,并提升系数大小的可解释性。

实验结果

研究问题

  • RQ1当预测变量高度相关时,是否存在一种变量选择方法能够保持符号一致性?
  • RQ2SPAC-Lasso在高度多重共线性条件下是否优于Lasso、自适应Lasso、SCAD和PC-simple等传统方法?
  • RQ3SPAC能否在部分相关性失效的相关预测变量中有效捕捉信号强度?
  • RQ4在标准正则性条件下,SPAC-Lasso是否在低维和高维设置下均具有一致性?
  • RQ5SPAC-Lasso在具有强连锁不平衡关系的现实基因组数据(如HapMap数据)中表现如何?

主要发现

  • 在满足正则性条件时,SPAC-Lasso在有限维和高维设置下均实现了符号一致性,确保了正确的变量选择。
  • 在高重共线性条件下,模拟研究显示该方法优于传统的Lasso、自适应Lasso、SCAD和PC-simple方法。
  • 在HapMap基因数据集中,SPAC-Lasso在预测变量因连锁不平衡而呈现强相关性的情况下,表现出更优的相关预测变量识别能力。
  • SPAC方法成功降低了相关性影响,同时保留了系数大小,从而更有效地检测到与响应变量的直接关联。
  • 理论分析表明,SPAC-Lasso即使在预测变量高度相关时仍能保持一致性,克服了标准惩罚方法的关键局限性。
  • 实证结果表明,SPAC-Lasso在变量选择准确性和信号检测方面优于基于部分相关性和标准惩罚方法的模型。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。