[论文解读] Asymptotic causal inference with observational studies trimmed by the estimated propensity scores
本文提出一种平滑加权方法,通过用连续可微的加权函数替代基于极端倾向得分的单位二值化截断,以解决观察性研究中重叠不足的问题。该方法确保渐近线性及有效的自助法推断,相较于传统截断方法提高了精度,同时基于明确定义的目标总体保持清晰的因果 estimand。
Causal inference with observational studies often relies on the assumptions of unconfoundedness and overlap of covariate distributions in different treatment groups. The overlap assumption is violated when some units have propensity scores close to 0 or 1, and therefore both practical and theoretical researchers suggest dropping units with extreme estimated propensity scores. However, existing trimming methods ignore the uncertainty in this design stage and restrict inference only to the trimmed sample, due to the non-smoothness of the trimming. We propose a smooth weighting, which approximates the existing sample trimming but has better asymptotic properties. An advantage of the new smoothly weighted estimator is its asymptotic linearity, which ensures that the bootstrap can be used to make inference for the target population, incorporating uncertainty arising from both the design and analysis stages. We also extend the theory to the average treatment effect on the treated, suggesting trimming samples with estimated propensity scores close to 1.
研究动机与目标
- 解决因果推断中样本截断的理论模糊性,其中截断会因样本不同而改变目标 estimand。
- 开发一种二值截断的平滑加权替代方法,保持渐近性质并支持有效推断。
- 将该框架扩展至处理组上的平均处理效应(ATT),基于高倾向得分进行截断。
- 确保新估计量渐近线性,并支持基于自助法的置信区间。
- 通过连续加权提高有限样本精度,避免指示函数的非光滑性。
提出的方法
- 基于排除极端估计倾向得分单位的目标总体定义因果参数,确保 estimand 稳定。
- 提出一种平滑权重函数,近似二值包含指示函数但具备可微性,减少非光滑性问题。
- 使用调优参数 ε 控制平滑度,当 ε → 0 时恢复原始截断规则。
- 使用标准两步估计量线性化技术,推导平滑加权估计量的渐近线性。
- 应用自助法构建平滑加权方案下因果效应的有效置信区间。
- 通过基于高倾向得分(如 0.78)截断单位,将该方法扩展至处理组上的平均处理效应(ATT)。
实验结果
研究问题
- RQ1如何重新定义因果推断中的样本截断,以确保目标总体和 estimand 的稳定与明确定义?
- RQ2当截断被平滑加权替代时,加权估计量的渐近性质是什么?
- RQ3能否可靠地使用自助法构建平滑加权估计量的置信区间?
- RQ4在有限样本中,平滑加权相比二值截断如何提升精度?
- RQ5在估计处理组上的平均处理效应(ATT)时,最优截断规则是什么?
主要发现
- 所提出的平滑加权估计量渐近线性,支持使用自助法构建有效置信区间。
- 平滑加权方法相比二值截断降低了方差,这一结果在渐近理论与模拟研究中均得到验证。
- 在国家支持工作示范项目中,平滑加权估计量对 ATT 的点估计为 1527,标准误为 397,而二值截断方法的结果为 1506 和 404。
- 平滑估计量的 95% 置信区间(697, 2314)比原始 Hainmueller(2012)的区间(97, 3044)更窄,与精度提升一致。
- 增广加权估计量未在精度上优于非增广版本,可能由于结果回归中的模型设定错误。
- 当 ε → 0 时,平滑权重函数收敛于指示函数,表明与现有截断实践的一致性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。