Skip to main content
QUICK REVIEW

[论文解读] A Graph-Theoretic Approach to Randomization Tests of Causal Effects Under General Interference

David Puelz, Guillaume Basse|arXiv (Cornell University)|Oct 23, 2019
Complex Network Analysis Techniques被引用 5
一句话总结

本文提出了一种图论框架,用于在一般干扰下进行有效的随机化检验,通过将因果零假设建模为二分图,并识别出双 clique(完全子图,其中零假设为精确值),从而实现条件随机化检验。该方法通过基于观测分配进行条件化来提升统计功效,并利用现成的图算法,在聚类和空间干扰设置下优于先前的方法,如在麦德林警务实验中所展示的。

ABSTRACT

Interference exists when a unit's outcome depends on another unit's treatment assignment. For example, intensive policing on one street could have a spillover effect on neighboring streets. Classical randomization tests typically break down in this setting because many null hypotheses of interest are no longer sharp under interference. A promising alternative is to instead construct a conditional randomization test on a subset of units and assignments for which a given null hypothesis is sharp. Finding these subsets is challenging, however, and existing methods are limited to special cases or have limited power. In this paper, we propose valid and easy-to-implement randomization tests for a general class of null hypotheses under arbitrary interference between units. Our key idea is to represent the hypothesis of interest as a bipartite graph between units and assignments, and to find an appropriate biclique of this graph. Importantly, the null hypothesis is sharp within this biclique, enabling conditional randomization-based tests. We also connect the size of the biclique to statistical power. Moreover, we can apply off-the-shelf graph clustering methods to find such bicliques efficiently and at scale. We illustrate our approach in settings with clustered interference and show advantages over methods designed specifically for that setting. We then apply our method to a large-scale policing experiment in Medellin, Colombia, where interference has a spatial structure.

研究动机与目标

  • 解决在因果推断中干扰违反无干扰假设时,进行有效随机化检验的挑战。
  • 克服现有方法的局限性,这些方法受限于特定的干扰结构(例如聚类或空间结构),缺乏普适性。
  • 开发一种构造性、模块化的框架,将统计功效与计算转化为零暴露图上的图论操作。
  • 通过基于观测分配进行条件化而非随机条件化,实现更高的统计功效。
  • 提供一种适用于任意干扰模式的一般性解决方案,包括复杂的空间和聚类结构。

提出的方法

  • 构建一个零暴露图作为二分图,其中单位和分配为节点,边表示在零假设下的观测到的零暴露。
  • 在零暴露图中,将双 clique 定义为完全二分子图,其中所有单位均与所有分配相连,确保在子集内零假设为精确值。
  • 使用双 clique 作为条件随机化检验的条件集,保证在零假设下检验的有效性。
  • 应用现成的图聚类算法,高效识别出较大的双 clique,以最大化统计功效。
  • 通过引入多零暴露图,将方法推广至交集零假设,以捕捉多个潜在结果约束。
  • 更新双 clique 查找过程,以保留潜在结果可推断性的分配为初始,适用于交集零假设。

实验结果

研究问题

  • RQ1当经典方法因非精确零假设而失效时,如何在一般干扰下构建有效的随机化检验?
  • RQ2我们能否识别出使给定零假设变为精确值的单位和分配子集,从而实现条件随机化检验?
  • RQ3所识别双 clique 的大小与在干扰下的随机化检验中统计功效有何关系?
  • RQ4该方法能否在无需事先结构假设的情况下,应用于复杂干扰结构(如空间或聚类干扰)?
  • RQ5在干扰设置下,基于观测分配进行条件化相比随机条件化,如何提升统计功效?

主要发现

  • 所提出的基于双 clique 的随机化检验在任意干扰下具有条件有效性,因为所识别双 clique 内部的零假设为精确值。
  • 该方法通过基于观测分配进行条件化,实现了更高的统计功效,这在模拟实验和麦德林警务实验中已得到验证。
  • 在麦德林实验中,原始结果的 FRT 显示所有半径下均存在显著溢出效应,但使用双 clique 方法调整结果后,未发现显著溢出效应,表明存在协变量混淆。
  • 调整结果的随机化分布方差更低,且中心位于比原始结果更小的正值区域,表明推断更精确。
  • 对调整结果进行的双 clique 检验显示 p 值与无显著溢出效应一致,与回归结果一致,后者在 0.05 水平下未发现显著系数。
  • 该方法在麦德林数据中成功识别出一个大型双 clique,使得在复杂空间干扰下实现了强大且有效的检验,优于针对聚类干扰设计的方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。