[Paper Review] Safe Policy Learning under Regression Discontinuity Designs with Multiple Cutoffs
This paper proposes a robust optimization framework for safe policy learning under regression discontinuity designs with multiple cutoffs, leveraging smooth heterogeneity across subpopulations to enable credible extrapolation. It ensures learned treatment cutoffs do not degrade overall utility compared to the status quo, achieving asymptotic regret bounds via doubly-robust estimation and partial identification under smoothness assumptions.
The regression discontinuity (RD) design is widely used for program evaluation with observational data. The primary focus of the existing literature has been the estimation of the local average treatment effect at the existing treatment cutoff. In contrast, we consider policy learning under the RD design. Because the treatment assignment mechanism is deterministic, learning better treatment cutoffs requires extrapolation. We develop a robust optimization approach to finding optimal treatment cutoffs that improve upon the existing ones. We first decompose the expected utility into point-identifiable and unidentifiable components. We then propose an efficient doubly-robust estimator for the identifiable parts. To account for the unidentifiable components, we leverage the existence of multiple cutoffs that are common under the RD design. Specifically, we assume that the heterogeneity in the conditional expectations of potential outcomes across different groups vary smoothly along the running variable. Under this assumption, we minimize the worst case utility loss relative to the status quo policy. The resulting new treatment cutoffs have a safety guarantee that they will not yield a worse overall outcome than the existing cutoffs. Finally, we establish the asymptotic regret bounds for the learned policy using semi-parametric efficiency theory. We apply the proposed methodology to empirical and simulated data sets.
Motivation & Objective
- To address the gap in policy learning under regression discontinuity (RD) designs, where existing methods focus on treatment effect estimation at fixed cutoffs.
- To develop a method for learning improved treatment cutoffs that extrapolate beyond the current policy while ensuring safety against worse outcomes.
- To exploit multiple subpopulation cutoffs—common in real-world RD applications—to enable credible extrapolation through smoothness assumptions on outcome heterogeneity.
- To provide theoretical safety guarantees that the new policy will not underperform the status quo, even under partial identification.
- To establish asymptotic regret bounds for the learned policy using semi-parametric efficiency theory.
Proposed method
- Decomposes expected utility into point-identified and partially-identified components using potential outcomes under different treatment rules.
- Proposes a doubly-robust estimator for the point-identified portion of utility, combining outcome regression and propensity score models.
- Imposes a smoothness assumption on cross-group differences in conditional potential outcomes across the running variable, with bounded slope in extrapolation regions.
- Uses robust optimization to minimize worst-case utility loss relative to the status quo policy, ensuring safety guarantees.
- Employs a model class for unidentifiable components based on bounds derived from the smoothness assumption, enabling worst-case utility evaluation.
- Derives asymptotic regret bounds using semi-parametric efficiency theory, showing convergence rates under regularity conditions.
Experimental results
Research questions
- RQ1Can we learn better treatment cutoffs in RD designs without assuming unconfoundedness or strong ignorability?
- RQ2How can we ensure that a new policy based on extrapolated cutoffs will not perform worse than the current policy?
- RQ3What role do multiple subpopulation cutoffs play in enabling credible extrapolation in RD designs?
- RQ4How can we estimate utility components that are partially identified due to extrapolation beyond observed data?
- RQ5What are the theoretical convergence rates (regret bounds) of the proposed safe policy learning method?
Key findings
- The proposed method ensures that the learned policy will not yield a worse overall utility than the status quo, providing a safety guarantee under robust optimization.
- The doubly-robust estimator achieves root-n consistency for the identifiable component of utility under regularity conditions.
- The worst-case utility loss is bounded by a term that decays at rate O_p(ρ_n^{-1}), where ρ_n is the sample size in the extrapolation region.
- Asymptotic regret bounds are established using semi-parametric efficiency theory, showing convergence to the optimal policy.
- The method achieves finite-sample safety by minimizing worst-case loss over a partially identified model class defined by smoothness constraints.
- Empirical application to Colombia’s ACCES loan program shows improved national enrollment outcomes using learned cutoffs per department.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.