Skip to main content
QUICK REVIEW

[论文解读] The Projected Covariance Measure for assumption-lean variable significance testing

Anton Rask Lundborg, Ilmun Kim|arXiv (Cornell University)|Nov 3, 2022
Statistical Methods and Inference被引用 9
一句话总结

本文提出了投影协方差度量(PCM),一种无需模型假设的条件均值独立性检验方法,利用灵活的非参数或机器学习方法(如样条和平滑回归树)进行检验。通过将数据分割以估计投影和条件协方差,PCM 实现了稳健的误差控制和极小化最大风险最优的检验功效,在模拟实验中优于参数检验和现有非参数检验方法。

ABSTRACT

Testing the significance of a variable or group of variables $X$ for predicting a response $Y$, given additional covariates $Z$, is a ubiquitous task in statistics. A simple but common approach is to specify a linear model, and then test whether the regression coefficient for $X$ is non-zero. However, when the model is misspecified, the test may have poor power, for example when $X$ is involved in complex interactions, or lead to many false rejections. In this work we study the problem of testing the model-free null of conditional mean independence, i.e. that the conditional mean of $Y$ given $X$ and $Z$ does not depend on $X$. We propose a simple and general framework that can leverage flexible nonparametric or machine learning methods, such as additive models or random forests, to yield both robust error control and high power. The procedure involves using these methods to perform regressions, first to estimate a form of projection of $Y$ on $X$ and $Z$ using one half of the data, and then to estimate the expected conditional covariance between this projection and $Y$ on the remaining half of the data. While the approach is general, we show that a version of our procedure using spline regression achieves what we show is the minimax optimal rate in this nonparametric testing problem. Numerical experiments demonstrate the effectiveness of our approach both in terms of maintaining Type I error control, and power, compared to several existing approaches.

研究动机与目标

  • 解决当模型误设导致检验功效低下或第一类错误膨胀时,参数模型在变量显著性检验中的局限性。
  • 开发一种无需模型假设的条件均值独立性检验方法,其中给定 X 和 Z 时 Y 的条件均值不依赖于 X。
  • 在保持有效第一类错误控制的前提下,使灵活的非参数或机器学习方法(如样条、随机森林)得以应用。
  • 在非参数条件均值独立性检验中实现极小化最大风险最优的检验速率。
  • 提供一种通用、假设宽松的框架,在模型误设下优于现有方法,在功效和误差控制方面表现更优。

提出的方法

  • 将数据分为两部分:第一部分用于使用非参数回归方法估计 Y 对 X 和 Z 的投影。
  • 第二部分用于估计投影后的 Y 与原始 Y 之间的条件协方差,构成检验统计量。
  • 检验统计量构造为在数据第二部分上计算投影响应与观测响应之间的经验协方差。
  • 通过样本分割确保估计与检验统计量之间的独立性,从而实现有效的推断。
  • 该方法具有通用性,可与任意回归方法结合使用,包括加法模型、样条和随机森林。
  • 理论分析表明,采用样条回归的 PCM 版本在非参数条件均值独立性检验中达到了极小化最大风险最优速率。

实验结果

研究问题

  • RQ1能否构建一种非参数、无需模型假设的条件均值独立性检验方法,在模型误设下仍能保持有效的第一类错误控制并实现高功效?
  • RQ2如何将灵活的机器学习方法整合进正式的假设检验框架中,而不会导致第一类错误膨胀?
  • RQ3所提出的投影协方差度量是否在非参数条件均值独立性检验中达到极小化最大风险最优速率?
  • RQ4在模型误设下,PCM 与现有方法(如 F 检验、GCM 和 wgsc)相比,在功效和误差控制方面表现如何?
  • RQ5PCM 是否能够检测到参数检验无法捕捉的复杂交互作用和异方差性?

主要发现

  • PCM 在多种数据生成机制下均保持有效的第一类错误控制,包括存在复杂交互作用和异方差性的情形。
  • 当采用样条回归时,PCM 实现了极小化最大风险最优收敛速率,确立了在非参数检验中的理论最优性。
  • 在模拟实验中,PCM 在功效方面优于传统的稳健标准误 F 检验和 wgsc 方法,尤其在模型误设条件下表现更优。
  • 当真实关系包含非线性关系和交互作用时,PCM 仍保持良好的功效,但在仅存在投影无法捕捉的纯交互效应的设定下可能损失功效。
  • wgsc 方法虽然有效,但在某些情形(尤其是二值响应)下表现出第一类错误膨胀,而 PCM 始终保持控制。
  • PCM 在不同样本量和回归方法(包括随机森林和加法模型)下均表现出稳健的性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。