[Paper Review] On the reliability of published findings using the regression discontinuity design in political science
This paper investigates the reliability of regression discontinuity (RD) findings in political science by analyzing 2009–2018 publications in top journals. It reveals that published RD estimates exhibit pathological bunching just above the 5% significance threshold, driven by underpowered studies and flawed inference methods—particularly underestimated standard errors—leading to exaggerated statistical significance despite minimal changes in point estimates upon reanalysis using modern methods.
The regression discontinuity (RD) design offers identification of causal effects under weak assumptions, earning it a position as a standard method in modern political science research. But identification does not necessarily imply that causal effects can be estimated accurately with limited data. In this paper, we highlight that estimation under the RD design involves serious statistical challenges and investigate how these challenges manifest themselves in the empirical literature in political science. We collect all RD-based findings published in top political science journals in the period 2009-2018. The distribution of published results exhibits pathological features; estimates tend to bunch just above the conventional level of statistical significance. A reanalysis of all studies with available data suggests that researcher discretion is not a major driver of these features. However, researchers tend to use inappropriate methods for inference, rendering standard errors artificially small. A retrospective power analysis reveals that most of these studies were underpowered to detect all but large effects. The issues we uncover, combined with well-documented selection pressures in academic publishing, cause concern that many published findings using the RD design may be exaggerated.
Motivation & Objective
- To assess the reliability of regression discontinuity (RD) findings published in top political science journals from 2009 to 2018.
- To investigate whether pathological patterns in reported t-statistics—bunching just above 1.96—stem from researcher discretion or methodological flaws.
- To evaluate whether inappropriate inference procedures, such as biased standard error estimators, contribute to overstatement of statistical significance.
- To examine the role of publication bias and low statistical power in distorting the empirical literature.
- To reanalyze published RD studies using state-of-the-art methods to assess the robustness of original findings.
Proposed method
- Collected all RD-based studies published in *American Political Science Review*, *American Journal of Political Science*, and *Journal of Politics* from 2009 to 2018.
- Conducted a retrospective power analysis to assess the statistical power of original studies to detect meaningful effects.
- Reanalyzed all studies with available data using standardized, modern methods (e.g., Calonico et al., 2015; rdrobust and rdhonest packages in R) to correct for methodological shortcomings.
- Compared original t-statistics with those from reanalysis using robust standard errors and bias-corrected inference procedures.
- Used funnel plots and p-value distributions to visualize the extent of selection bias and overestimation of significance.
- Evaluated bandwidth selection methods (automated vs. non-automated) to test whether researcher discretion explains the observed bunching.
Experimental results
Research questions
- RQ1Do published RD estimates in political science exhibit pathological bunching just above the conventional 5% significance threshold?
- RQ2To what extent is researcher discretion in bandwidth selection responsible for the observed bunching in t-statistics?
- RQ3How do modern inference methods affect the standard errors and statistical significance of RD estimates compared to original analyses?
- RQ4What is the statistical power of RD studies in political science, and how does this relate to the prevalence of false positives?
- RQ5To what extent do publication bias and underpowered studies contribute to the overstatement of effect sizes in the literature?
Key findings
- The distribution of reported t-statistics shows significant bunching just above 1.96, indicating a systematic overrepresentation of results just meeting conventional significance thresholds.
- Reanalysis using modern methods (rdrobust and rdhonest) increased standard errors on average, resulting in a leftward shift of t-statistics toward zero, suggesting original findings were overstated in terms of statistical significance.
- Studies using automated bandwidth selection showed more bunching than non-automated ones, indicating that researcher discretion in bandwidth choice is not the primary driver of the pathological pattern.
- Retrospective power analysis revealed that most studies were underpowered to detect all but large effects, increasing the risk of false positives.
- The combination of low power and flawed inference procedures—especially biased standard error estimators—exacerbates the risk of false positives and inflates the proportion of false discoveries in the literature.
- Even after correcting for methodological flaws, point estimates remained largely unchanged, but their precision was substantially reduced, indicating that the original findings were overconfident in their significance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.