Skip to main content
QUICK REVIEW

[论文解读] Safe Policy Learning under Regression Discontinuity Designs with Multiple Cutoffs

Yi Zhang, Eli Ben‐Michael|arXiv (Cornell University)|Aug 29, 2022
Advanced Causal Inference Techniques被引用 4
一句话总结

本文提出了一种在具有多个临界值的回归不连续设计下,针对安全策略学习的鲁棒优化框架,通过利用子群体间结果异质性的平滑性,实现可信的外推。该方法确保所学习的处理临界值不会使整体效用低于现状,通过双重稳健估计和在平滑性假设下的部分识别,实现渐近遗憾界。

ABSTRACT

The regression discontinuity (RD) design is widely used for program evaluation with observational data. The primary focus of the existing literature has been the estimation of the local average treatment effect at the existing treatment cutoff. In contrast, we consider policy learning under the RD design. Because the treatment assignment mechanism is deterministic, learning better treatment cutoffs requires extrapolation. We develop a robust optimization approach to finding optimal treatment cutoffs that improve upon the existing ones. We first decompose the expected utility into point-identifiable and unidentifiable components. We then propose an efficient doubly-robust estimator for the identifiable parts. To account for the unidentifiable components, we leverage the existence of multiple cutoffs that are common under the RD design. Specifically, we assume that the heterogeneity in the conditional expectations of potential outcomes across different groups vary smoothly along the running variable. Under this assumption, we minimize the worst case utility loss relative to the status quo policy. The resulting new treatment cutoffs have a safety guarantee that they will not yield a worse overall outcome than the existing cutoffs. Finally, we establish the asymptotic regret bounds for the learned policy using semi-parametric efficiency theory. We apply the proposed methodology to empirical and simulated data sets.

研究动机与目标

  • 为解决现有方法在回归不连续(RD)设计下仅关注固定临界值处处理效应估计的局限性,填补政策学习中的空白。
  • 开发一种学习改进处理临界值的方法,使其能够外推至当前政策之外,同时确保不会导致更差的结果。
  • 利用现实世界RD应用中常见的多个子群体临界值,通过结果异质性平滑性假设,实现可信的外推。
  • 提供理论上的安全保证,确保新政策即使在部分识别条件下也不会劣于现状。
  • 基于半参数效率理论,建立所学习策略的渐近遗憾界。

提出的方法

  • 通过不同处理规则下的潜在结果,将期望效用分解为点识别和部分识别两部分。
  • 为效用的点识别部分提出双重稳健估计量,结合结果回归模型与倾向得分模型。
  • 对运行变量上跨组条件潜在结果差异施加平滑性假设,外推区域内的斜率有界。
  • 使用鲁棒优化方法,最小化相对于现状政策的最坏情况效用损失,从而确保安全保证。
  • 基于平滑性假设推导的边界,构建不可识别成分的模型类,以支持最坏情况效用评估。
  • 利用半参数效率理论推导渐近遗憾界,表明在正则条件下具有收敛速率。

实验结果

研究问题

  • RQ1我们能否在不假设无混淆性或强可忽略性的情况下,学习到RD设计中更优的处理临界值?
  • RQ2如何确保基于外推临界值的新政策不会表现得比当前政策更差?
  • RQ3多个子群体临界值在RD设计中如何促进可信外推?
  • RQ4如何估计因外推至观测数据范围之外而部分识别的效用分量?
  • RQ5所提出的可靠策略学习方法的理论收敛速率(遗憾界)是什么?

主要发现

  • 所提方法确保所学习的策略不会导致整体效用低于现状,通过鲁棒优化提供安全保证。
  • 在正则条件下,双重稳健估计量对效用的可识别分量实现根n一致性。
  • 最坏情况效用损失被一个以速率O_p(ρ_n^{-1})衰减的项所界定,其中ρ_n为外推区域的样本量。
  • 基于半参数效率理论,建立了渐近遗憾界,表明策略收敛至最优策略。
  • 通过在由平滑性约束定义的部分识别模型类上最小化最坏情况损失,实现有限样本下的安全性。
  • 在哥伦比亚ACCES贷款计划中的实证应用表明,使用按部门学习的临界值可提升全国入学率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。