Skip to main content
QUICK REVIEW

[论文解读] A paradox from randomization-based causal inference

Peng Ding|arXiv (Cornell University)|Feb 2, 2014
Advanced Causal Inference Techniques参考文献 25被引用 12
一句话总结

本文揭示了基于随机化的因果推断中一个反直觉的悖论:尽管费舍尔的零假设(个体因果效应为零)在逻辑上蕴含内曼的零假设(平均因果效应为零),但在实践中,内曼的检验可能拒绝原假设,而费舍尔的检验却无法拒绝——尤其在恒定因果效应下更为明显。作者通过渐近分析、模拟实验和真实数据,在完全随机化、分层、配对设计及因子设计实验中验证了这一悖论,揭示了尽管存在逻辑蕴含关系,两种检验行为之间仍存在根本性差异。

ABSTRACT

Under the potential outcomes framework, causal effects are defined as comparisons between potential outcomes under treatment and control. To infer causal effects from randomized experiments, Neyman proposed to test the null hypothesis of zero average causal effect (Neyman's null), and Fisher proposed to test the null hypothesis of zero individual causal effect (Fisher's null). Although the subtle difference between Neyman's null and Fisher's null has caused lots of controversies and confusions for both theoretical and practical statisticians, a careful comparison between the two approaches has been lacking in the literature for more than eighty years. We fill in this historical gap by making a theoretical comparison between them and highlighting an intriguing paradox that has not been recognized by previous researchers. Logically, Fisher's null implies Neyman's null. It is therefore surprising that, in actual completely randomized experiments, rejection of Neyman's null does not imply rejection of Fisher's null for many realistic situations, including the case with constant causal effect. Furthermore, we show that this paradox also exists in other commonly-used experiments, such as stratified experiments, matched-pair experiments, and factorial experiments. Asymptotic analyses, numerical examples, and real data examples all support this surprising phenomenon. Besides its historical and theoretical importance, this paradox also leads to useful practical implications for modern researchers.

研究动机与目标

  • 解决内曼与费舍尔基于随机化的因果推断框架之间长期存在的理论模糊性。
  • 探究为何在逻辑蕴含关系下,内曼检验对零平均因果效应的拒绝能力可能强于费舍尔检验对零个体效应的检验能力。
  • 证明该悖论并非人为产物,而是常见实验设计中的系统性现象。
  • 为研究人员在随机化实验中选择检验方法与解释结果提供实用指导。

提出的方法

  • 在潜在结果框架下,对内曼的零假设(零平均因果效应)与费舍尔的零假设(零个体因果效应)进行理论比较。
  • 对两种原假设下的检验功效进行渐近分析,表明当因果效应恒定时,内曼检验的功效更高。
  • 通过数值模拟与来自$2^4$因子设计的真实数据示例,展示该悖论在实践中的表现。
  • 采用随机化推断,包括费舍尔随机化检验(FRT)与得分检验,评估在不同原假设下的拒绝率。
  • 推导出该悖论发生的条件,基于潜在结果之间的相关性与处理分配概率。
  • 将结果扩展至分层、配对与因子设计实验,以证明该悖论的普遍性。

实验结果

研究问题

  • RQ1在何种条件下,尽管费舍尔的零假设在逻辑上蕴含内曼的零假设,内曼检验仍能拒绝原假设,而费舍尔检验却无法拒绝?
  • RQ2为何在恒定因果效应假设下,内曼检验的功效高于费舍尔检验,这与逻辑蕴含关系的预期相悖?
  • RQ3该悖论是否在完全随机化实验之外的实验设计中依然存在,如分层或因子设计?
  • RQ4不同检验统计量(如t统计量、Kolmogorov–Smirnov检验、Wilcoxon–Mann–Whitney检验)如何影响该悖论的行为?
  • RQ5该悖论对研究人员在随机化实验中选择内曼式或费舍尔式推断具有何种实际影响?

主要发现

  • 在完全随机化实验中,当因果效应恒定时,内曼检验可能拒绝原假设,而费舍尔检验却无法拒绝。
  • 该悖论的根源在于,当因果效应恒定时,内曼检验的渐近功效高于费舍尔检验,尽管费舍尔的零假设逻辑上蕴含内曼的零假设。
  • 该悖论不仅存在于完全随机化实验中,也存在于分层、配对及因子设计实验中,该结论得到渐近理论与真实数据的双重验证。
  • 在费舍尔的精确零假设下,Kolmogorov–Smirnov与Wilcoxon–Mann–Whitney检验统计量的随机化分布比在内曼的平均零假设下更分散,导致在平均零假设下FRT检验过于保守。
  • 该悖论发生的条件取决于潜在结果与处理分配概率之间的相关性,临界阈值位于黄金分割比例的倒数(≈0.618)。
  • 该悖论有助于解释为何在因子设计实验中,基于反转FRT的区间估计量通常比内曼式置信区间更宽,如Dasgupta等(2015)所观察到的现象。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。