Skip to main content
QUICK REVIEW

[论文解读] The conditional permutation test

Thomas B. Berrett, Yi Wang|arXiv (Cornell University)|Jul 14, 2018
Bayesian Methods and Mixture Models被引用 6
一句话总结

本文提出了条件置换检验(conditional permutation test),一种用于在高维混杂因子 Z 的条件下评估 X 与 Y 条件独立性的方法。该方法通过非均匀地置换 X 的取值,同时保持其与 Z 的依赖关系,实现对条件独立性的检验。该方法利用估计的条件分布 X|Z 来指导置换权重,即使在存在近似误差的情况下,也能确保有效的 I 类错误控制,且其错误边界与条件随机化检验(conditional randomization test)相当。

ABSTRACT

We propose a general new method, the \emph{conditional permutation test}, for testing the conditional independence of variables $X$ and $Y$ given a potentially high-dimensional random vector $Z$ that may contain confounding factors. The proposed test permutes entries of $X$ non-uniformly, so as to respect the existing dependence between $X$ and $Z$ and thus account for the presence of these confounders. Like the conditional randomization test of \citet{candes2018panning}, our test relies on the availability of an approximation to the distribution of $X \mid Z$---while \citet{candes2018panning}'s test uses this estimate to draw new $X$ values, for our test we use this approximation to design an appropriate non-uniform distribution on permutations of the $X$ values already seen in the true data. We provide an efficient Markov Chain Monte Carlo sampler for the implementation of our method, and establish bounds on the Type~I error in terms of the error in the approximation of the conditional distribution of $X\mid Z$, finding that, for the worst case test statistic, the inflation in Type I error of the conditional permutation test is no larger than that of the conditional randomization test. We validate these theoretical results with experiments on simulated data and on the Capital Bikeshare data set.

研究动机与目标

  • 解决在存在可能扭曲标准独立性检验的高维混杂因子 Z 的情况下,测试条件独立性所面临的挑战。
  • 开发一种在置换过程中尊重 X 与 Z 之间依赖结构的方法,避免因混杂导致的无效推断。
  • 确保在条件分布 X|Z 为估计值而非精确已知时,仍能实现 I 类错误控制。
  • 提供一种计算效率较高的替代方法,相较于条件随机化检验,具有相似的误差保证。
  • 在模拟数据和真实的 Capital Bikeshare 数据集上对方法进行实证验证。

提出的方法

  • 提出一种针对 X 的非均匀置换方案,通过使用估计的条件分布 X|Z 来保留 X 与 Z 之间观察到的依赖关系。
  • 设计一种马尔可夫链蒙特卡洛(MCMC)采样器,以高效地从由 X|Z 近似定义的非均匀置换分布中抽样。
  • 利用估计的 X|Z 为每种置换后的 X 配置分配置换概率,以反映其在条件模型下的合理性。
  • 建立 I 类错误膨胀的理论边界,表明其最大值不超过在相同近似误差下条件随机化检验的误差。
  • 将该检验方法应用于模拟数据和 Capital Bikeshare 数据集,以验证其经验性能和鲁棒性。

实验结果

研究问题

  • RQ1当混杂因子 Z 为高维且仅近似已知时,基于置换的检验能否维持有效的 I 类错误?
  • RQ2基于 X|Z 估计的非均匀置换相较于均匀置换,在条件独立性检验中如何改进?
  • RQ3X|Z 的近似误差与条件置换检验中 I 类错误膨胀之间的理论关系是什么?
  • RQ4与条件随机化检验相比,条件置换检验在性能和误差控制方面表现如何?
  • RQ5所提出的方法能否高效实现并应用于具有复杂混杂结构的真实世界数据集?

主要发现

  • 条件置换检验在 X|Z 条件分布被近似时,仍能维持 I 类错误控制,且错误膨胀的边界与条件随机化检验相当。
  • 对于最坏情况下的检验统计量,当使用相同的 X|Z 近似时,条件置换检验的 I 类错误最大膨胀程度不超过条件随机化检验。
  • MCMC 采样器实现了非均匀置换的高效实现,使该方法可扩展至高维 Z。
  • 在模拟数据上的实证验证表明,理论错误边界在实践中成立。
  • 在 Capital Bikeshare 数据集上的应用展示了该方法在具有复杂混杂因素的真实场景中的实用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。