Skip to main content
QUICK REVIEW

[论文解读] A theoretical treatment of conditional independence testing under Model-X

Eugene Katsevich, Aaditya Ramdas|arXiv (Cornell University)|May 12, 2020
Statistical Methods and Inference参考文献 35被引用 10
一句话总结

本文在Model-X (MX) 框架下建立了条件独立性检验的理论基础,通过Neyman-Pearson引理推导出最优检验统计量,并在最小矩条件假设下展示了统一的渐近第一类错误控制。通过在局部渐近正态性下推导统计功效表达式,将MX与经典统计检验及因果推断联系起来,实现了MX设定下的估计。

ABSTRACT

For testing conditional independence (CI) of a response $Y$ and a predictor $X$ given covariates $Z$, the recently introduced model-X (MX) framework has been the subject of active methodological research, especially in the context of MX knockoffs and their successful application to genome-wide association studies. In this paper, we build a theoretical foundation for the MX CI problem, yielding quantitative explanations for empirically observed phenomena and novel insights to guide the design of MX methodology. We focus our analysis on the conditional randomization test (CRT), whose validity conditional on $Y,Z$ allows us to view it as a test of a point null hypothesis involving the conditional distribution of $X$. We use the Neyman-Pearson lemma to derive the most powerful CRT statistic against a point alternative as well as an analogous result for MX knockoffs. We define CRT-style analogs of $t$- and $F$-tests with explicit critical values, and show that they have uniform asymptotic Type-I error control under the assumption that only the first two moments of $X$ given $Z$ are known, a significant relaxation of MX. We derive expressions for the power of these tests against local semiparametric alternatives using Le Cam's local asymptotic normality theory, explicitly capturing the prediction error of the underlying learning algorithm. Finally, we pave the way for estimation in the MX setting by drawing connections to semiparametric statistics and causal inference. Thus, this work forms explicit bridges from MX to both classical statistics (testing) and modern causal inference (estimation).

研究动机与目标

  • 在Model-X (MX) 框架下建立条件独立性检验的严格理论框架。
  • 通过正式的统计分析解释MX方法中经验观察到的行为。
  • 通过将MX方法与经典假设检验及半参数估计相联系,拓展MX方法论。
  • 在对协变量分布的建模假设最小的前提下,推导条件独立性的最优检验统计量。
  • 通过与半参数统计及因果推断的联系,实现在MX设定下的估计。

提出的方法

  • 利用Neyman-Pearson引理,在点备择假设下推导出最有力的条件随机化检验(CRT)统计量。
  • 在仅知$X$给$Z$的一阶和二阶矩的假设下,构建具有显式临界值的CRT风格$t$-和$F$-检验类比。
  • 应用Le Cam的局部渐近正态性(LAN)理论,推导对局部半参数替代假设的检验功效。
  • 通过显式功效表达式量化学习算法预测误差对MX检验功效的影响。
  • 在估计与因果推断背景下,建立MX与半参数统计之间的理论联系,特别是与半参数统计的联系。
  • 通过将CRT框架扩展至敲扑构造设定,推导出最优的MX敲扑统计量。

实验结果

研究问题

  • RQ1在MX框架下,给定点备择假设时,条件独立性检验的最有力统计量是什么?
  • RQ2在MX设定下,如何在仅基于$X|Z$的一阶和二阶矩最小假设下,构建具有有效临界值的$t$-和$F$-检验类比?
  • RQ3当仅知$X|Z$的一阶和二阶矩时,MX检验的渐近第一类错误控制行为如何?
  • RQ4条件分布估计器的预测误差如何影响MX检验的功效?
  • RQ5MX方法与经典半参数统计及因果推断框架之间存在何种联系?

主要发现

  • 通过Neyman-Pearson引理推导出在点备择假设下的最有力CRT统计量,为MX检验提供了理论基准。
  • 构建了具有显式临界值的CRT风格$t$-和$F$-检验类比,确保在仅对$X|Z$的一阶和二阶矩作假设时,实现统一的渐近第一类错误控制。
  • 利用Le Cam的LAN理论推导出MX检验对局部半参数替代假设的检验功效,显式纳入了条件均值估计器的预测误差。
  • 推导出的功效表达式表明,当$X$给$Z$的条件均值模型中预测误差增大时,检验性能会下降。
  • 本文建立了MX与半参数统计之间的理论联系,为未来在MX框架下开展估计研究奠定了基础。
  • 该框架为设计具有保证统计性质的MX敲扑及其他基于MX的推断程序提供了原则性基础。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。