[论文解读] A Power Analysis of the Conditional Randomization Test and Knockoffs
本文针对高维线性模型中 $ p/n \to \kappa > 0 $ 的情形,对条件随机化检验(CRT)和模型-X knockoffs进行了严格的渐近功效分析,推导了在各种检验统计量(边际协方差、OLS、Lasso)下的功效显式表达式。结果表明,在 i.i.d. 高斯协变量下,采用 Benjamini-Hochberg 或 AdaPT 方法计算 $ p $-值的 CRT 在功效上优于 knockoffs,并建立了利用未标记数据和回顾性抽样所带来的理论增益。
In many scientific problems, researchers try to relate a response variable $Y$ to a set of potential explanatory variables $X = (X_1,\dots,X_p)$, and start by trying to identify variables that contribute to this relationship. In statistical terms, this goal can be posed as trying to identify $X_j$'s upon which $Y$ is conditionally dependent. Sometimes it is of value to simultaneously test for each $j$, which is more commonly known as variable selection. The conditional randomization test (CRT) and model-X knockoffs are two recently proposed methods that respectively perform conditional independence testing and variable selection by, for each $X_j$, computing any test statistic on the data and assessing that test statistic's significance by comparing it to test statistics computed on synthetic variables generated using knowledge of $X$'s distribution. Our main contribution is to analyze their power in a high-dimensional linear model where the ratio of the dimension $p$ and the sample size $n$ converge to a positive constant. We give explicit expressions of the asymptotic power of the CRT, variable selection with CRT $p$-values, and model-X knockoffs, each with a test statistic based on either the marginal covariance, the least squares coefficient, or the lasso. One useful application of our analysis is the direct theoretical comparison of the asymptotic powers of variable selection with CRT $p$-values and model-X knockoffs; in the instances with independent covariates that we consider, the CRT provably dominates knockoffs. We also analyze the power gain from using unlabeled data in the CRT when limited knowledge of $X$'s distribution is available, and the power of the CRT when samples are collected retrospectively.
研究动机与目标
- 理论分析条件随机化检验(CRT)和模型-X knockoffs 在 $ p/n \to \kappa > 0 $ 的高维设定下的渐近功效。
- 比较基于 Benjamini-Hochberg 和 AdaPT 的 CRT 变量选择方法与模型-X knockoffs 在相同检验统计量下的功效。
- 量化在仅部分掌握 $ X $ 分布知识时,将未标记数据引入 CRT 所带来的功效增益。
- 将分析扩展至回顾性抽样情形,推导有效信号强度作为 $ Y $ 的边际二阶矩的函数。
- 通过有限样本模拟验证理论发现,显示与理论方法结果高度一致。
提出的方法
- 在协变量为多元高斯分布且 $ \Sigma = I $ 的条件下,利用三种检验统计量(边际协方差、OLS 系数、Lasso)推导 CRT 和 knockoffs 的渐近功效表达式。
- 应用近似消息传递(AMP)框架,刻画在高维极限下检验统计量的渐近行为。
- 通过 CDF 变换的 knockoff 统计量,证明 knockoffs 与应用于 CRT $ p $-值的 AdaPT 在渐近意义上等价。
- 通过建模具有已知矩的合成 $ X $-变量,将未标记数据纳入分析,推导出功效提升的边界。
- 通过将有效信号强度建模为 $ \mathbb{E}[Y^2] $ 的函数,将结果扩展至回顾性抽样情形。
- 运用随机矩阵理论和 Stein 引理中的理论工具,推导出信号强度和检验统计量分布的不动点方程。
实验结果
研究问题
- RQ1在高维线性模型中,当使用边际协方差、OLS 或 Lasso 作为检验统计量时,CRT 的渐近功效是多少?
- RQ2在相同的检验统计量和 i.i.d. 高斯协变量下,基于 Benjamini-Hochberg 或 AdaPT 的 CRT 变量选择功效与模型-X knockoffs 相比如何?
- RQ3在仅部分掌握 $ X $ 分布知识时,使用未标记数据对 CRT 的功效提升具有何种理论意义?
- RQ4回顾性抽样如何影响 CRT 的有效信号强度和功效?
- RQ5为 CRT 和 knockoffs 推导出的渐近功效表达式是否能在有限样本中得到准确验证?
主要发现
- 在 i.i.d. 高斯协变量($ \Sigma = I $)下,采用 Benjamini-Hochberg 或 AdaPT 方法处理边际协方差、OLS 或 Lasso 统计量的 $ p $-值时,CRT 在渐近意义上能将 FDR 控制在名义水平。
- 当 $ \Sigma = I $ 时,对任意三种检验统计量的 $ p $-值应用 AdaPT 的 CRT,在渐近功效上严格优于模型-X knockoffs。
- CRT 使用边际协方差检验统计量的功效随未标记样本数量增加而提升,且在有限未标记数据条件下推导出该增益的边界。
- 对于回顾性抽样数据,CRT 的有效信号强度被显式表征为 $ Y $ 的边际二阶矩的函数,从而实现功效的量化。
- 通过 AMP 框架推导出的理论功效表达式与有限样本模拟结果高度一致,表明 CRT 和 knockoffs 可实现接近理论最优方法的功效。
- knockoffs 在渐近意义上等价于对 CDF 变换后的 knockoff 统计量应用 AdaPT,从而实现了与基于 CRT 的方法的直接比较。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。