Skip to main content
QUICK REVIEW

[论文解读] On the reliability of published findings using the regression discontinuity design in political science

Drew Stommes, Peter M. Aronow|arXiv (Cornell University)|Sep 29, 2021
Qualitative Comparative Analysis Research被引用 4
一句话总结

本文通过分析2009–2018年顶级政治科学期刊中的发表文献,调查了政治科学中回归不连续设计(RD)结果的可靠性。研究发现,已发表的RD估计值在5%显著性阈值附近表现出病态的聚集现象,其成因是统计功效不足以及推断方法存在缺陷——尤其是低估了标准误——导致尽管点估计值在使用现代方法重新分析后变化极小,但统计显著性却被过度夸大。

ABSTRACT

The regression discontinuity (RD) design offers identification of causal effects under weak assumptions, earning it a position as a standard method in modern political science research. But identification does not necessarily imply that causal effects can be estimated accurately with limited data. In this paper, we highlight that estimation under the RD design involves serious statistical challenges and investigate how these challenges manifest themselves in the empirical literature in political science. We collect all RD-based findings published in top political science journals in the period 2009-2018. The distribution of published results exhibits pathological features; estimates tend to bunch just above the conventional level of statistical significance. A reanalysis of all studies with available data suggests that researcher discretion is not a major driver of these features. However, researchers tend to use inappropriate methods for inference, rendering standard errors artificially small. A retrospective power analysis reveals that most of these studies were underpowered to detect all but large effects. The issues we uncover, combined with well-documented selection pressures in academic publishing, cause concern that many published findings using the RD design may be exaggerated.

研究动机与目标

  • 评估2009至2018年间发表于顶级政治科学期刊的回归不连续(RD)研究结果的可靠性。
  • 调查报告的t统计量中出现的病态模式(即在1.96附近聚集)是否源于研究者自由裁量权,还是方法论缺陷所致。
  • 评估不恰当的推断程序(如存在偏倚的标准误估计器)是否导致对统计显著性的夸大。
  • 考察发表偏倚和统计功效低下在多大程度上扭曲了实证文献。
  • 使用最先进的方法重新分析已发表的RD研究,以评估原始发现的稳健性。

提出的方法

  • 收集了2009至2018年间发表于《美国政治科学评论》(*American Political Science Review*)、《美国政治科学杂志》(*American Journal of Political Science*)和《政治学杂志》(*Journal of Politics*)的所有基于RD的研究。
  • 开展回顾性统计功效分析,评估原始研究检测有意义效应的能力。
  • 对所有可获得数据的研究,使用标准化的现代方法(如Calonico等人,2015年;R中的rdrobust和rdhonest包)进行重新分析,以纠正方法论缺陷。
  • 将原始t统计量与使用稳健标准误和偏差校正推断程序的重新分析结果进行比较。
  • 使用漏斗图和p值分布图,可视化选择偏倚和显著性被高估的程度。
  • 评估带宽选择方法(自动化与非自动化)的差异,以检验研究者自由裁量权是否解释了观察到的聚集现象。

实验结果

研究问题

  • RQ1政治科学中已发表的RD估计值是否在传统的5%显著性阈值附近表现出病态聚集?
  • RQ2研究者在带宽选择上的自由裁量权在多大程度上导致了t统计量的聚集现象?
  • RQ3与原始分析相比,现代推断方法如何影响RD估计的标准误和统计显著性?
  • RQ4政治科学中RD研究的统计功效如何?这与假阳性结果的普遍性有何关联?
  • RQ5发表偏倚和统计功效低下在多大程度上导致了文献中效应量的夸大?

主要发现

  • 报告的t统计量分布显示在1.96附近存在显著的聚集现象,表明在常规显著性阈值附近存在系统性的结果过度代表。
  • 使用现代方法(rdrobust和rdhonest)重新分析后,标准误平均增加,导致t统计量向零方向显著左移,表明原始发现的统计显著性被过度夸大。
  • 采用自动化带宽选择的研究比非自动化研究表现出更多的聚集现象,表明研究者在带宽选择上的自由裁量权并非导致该病态模式的主要原因。
  • 回顾性功效分析显示,大多数研究的统计功效不足以检测除大效应外的其他效应,从而增加了假阳性的风险。
  • 统计功效低下与推断方法缺陷(尤其是存在偏倚的标准误估计器)相结合,加剧了假阳性的风险,并提高了文献中错误发现的比例。
  • 即使在纠正了方法论缺陷后,点估计值基本保持不变,但其精度显著降低,表明原始发现对其显著性的判断过于自信。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。