Skip to main content
QUICK REVIEW

[论文解读] A Power Analysis for Knockoffs with the Lasso Coefficient-Difference Statistic

Asaf Weinstein, Weijie Su|arXiv (Cornell University)|Jul 30, 2020
Random Matrices and Applications参考文献 24被引用 17
一句话总结

本文提出了一种基于 knockoffs 校准的阈值 Lasso 方法,用于在高维线性模型中进行变量选择,利用近似消息传递(AMP)理论在无需先验信号知识的情况下控制错误发现率(FDR)。结果表明,当通过在扩展数据矩阵(含 knockoff 变量)上进行交叉验证选择 $\lambda$ 时,该方法在真正例比例方面显著优于标准 Lasso。

ABSTRACT

In a linear model with possibly many predictors, we consider variable selection procedures given by $$ \{1\leq j\leq p: |\widehat{\beta}_j(\lambda)| > t\}, $$ where $\widehat{\beta}(\lambda)$ is the Lasso estimate of the regression coefficients, and where $\lambda$ and $t$ may be data dependent. Ordinary Lasso selection is captured by using $t=0$, thus allowing to control only $\lambda$, whereas thresholded-Lasso selection allows to control both $\lambda$ and $t$. The potential advantages of the latter over the former in terms of power---figuratively, opening up the possibility to look further down the Lasso path---have been quantified recently leveraging advances in approximate message-passing (AMP) theory, but the implications are actionable only when assuming substantial knowledge of the underlying signal. In this work we study theoretically the power of a knockoffs-calibrated counterpart of thresholded-Lasso that enables us to control FDR in the realistic situation where no prior information about the signal is available. Although the basic AMP framework remains the same, our analysis requires a significant technical extension of existing theory in order to handle the pairing between original variables and their knockoffs. Relying on this extension we obtain exact asymptotic predictions for the true positive proportion achievable at a prescribed type I error level. In particular, we show that the knockoffs version of thresholded-Lasso can perform much better than ordinary Lasso selection if $\lambda$ is chosen by cross-validation on the augmented matrix.

研究动机与目标

  • 开发一种变量选择方法,可在高维线性模型中控制错误发现率(FDR),且无需依赖对潜在信号的先验知识。
  • 将近似消息传递(AMP)理论扩展至处理原始变量与其 knockoff 对应变量之间的配对关系。
  • 在固定 I 类错误水平下,量化基于 knockoff 的阈值 Lasso 的统计功效,以真正例比例为度量。
  • 证明在扩展数据矩阵(原始变量 + knockoff 变量)上进行交叉验证,可获得优于标准 Lasso 的性能。

提出的方法

  • 该方法采用阈值 Lasso 选择规则:若变量的 Lasso 回归系数绝对值超过一个依赖数据的阈值 $t$,则被选中,其中正则化参数 $\lambda$ 通过交叉验证进行调优。
  • 提出一种 knockoffs 校准程序,通过构建与原始预测变量配对的合成变量,确保 FDR 控制。
  • 分析基于近似消息传递(AMP)理论的扩展版本,用于在高维渐近条件下建模原始变量与 knockoff 变量的联合行为。
  • 关键统计量为原始变量与 knockoff 变量之间系数的差异,用于对特征进行排序与选择,同时控制 FDR。
  • 该方法作用于通过将 knockoff 变量附加到原始预测变量矩阵上形成的扩展设计矩阵。
  • 在随机设计和 i.i.d. 高斯噪声条件下,利用扩展的 AMP 框架推导出真正例比例的渐近预测。

实验结果

研究问题

  • RQ1在高维设置下,基于 knockoff 校准的阈值 Lasso 选择是否能实现比标准 Lasso 更高的统计功效?
  • RQ2当 $\lambda$ 的选择方式(特别是通过在扩展矩阵上进行交叉验证)不同时,其对真正例比例和 FDR 控制有何影响?
  • RQ3在固定 I 类错误率下,基于 knockoff 的阈值 Lasso 的渐近性能(以真正例比例衡量)如何?
  • RQ4原始变量与 knockoff 变量之间的配对结构如何影响变量选择方法的理论分析?
  • RQ5在不掌握信号结构先验知识的前提下,基于 knockoff 的选择方法的统计功效最多可提升多少?

主要发现

  • 当通过在扩展数据矩阵上进行交叉验证选择 $\lambda$ 时,knockoffs 校准的阈值 Lasso 方法在真正例比例方面显著优于普通 Lasso。
  • 该方法的渐近性能通过扩展的近似消息传递(AMP)框架得到理论刻画,该框架考虑了原始变量与 knockoff 变量的配对关系。
  • 在高维渐近条件下,推导出真正例比例的精确渐近预测,为 FDR 控制的变量选择提供了理论基准。
  • 该方法无需事先了解信号结构即可实现 FDR 控制,使其在真实场景中具有广泛适用性。
  • 在扩展矩阵上进行交叉验证显著提升了统计功效,证明了在正则化调优过程中利用 knockoff 信息的优势。
  • 理论框架证实,通过 knockoff 校准的阈值 Lasso 可通过沿 Lasso 路径进一步探索而优于标准 Lasso。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。