Skip to main content
QUICK REVIEW

[论文解读] Robust Estimation of Causal Effects via High-Dimensional Covariate Balancing Propensity Score

Yang Ning, Sida Peng|arXiv (Cornell University)|Dec 20, 2018
Advanced Causal Inference Techniques参考文献 39被引用 16
一句话总结

本文提出了高维协变量平衡倾向得分(HD-CBPS),一种在高维混杂因素的观察性研究中估计平均处理效应的稳健方法。通过将惩罚M估计与目标协变量平衡相结合,HD-CBPS即使在模型误设情况下也能实现根n一致性与渐近正态性,确保双重稳健性,并在无需变量选择一致性的情况下实现有效推断。

ABSTRACT

In this paper, we propose a robust method to estimate the average treatment effects in observational studies when the number of potential confounders is possibly much greater than the sample size. We first use a class of penalized M-estimators for the propensity score and outcome models. We then calibrate the initial estimate of the propensity score by balancing a carefully selected subset of covariates that are predictive of the outcome. Finally, the estimated propensity score is used to construct the inverse probability weighting estimator. We prove that the proposed estimator, which has the sample boundedness property, is root-n consistent, asymptotically normal, and semiparametrically efficient when the propensity score model is correctly specified and the outcome model is linear in covariates. More importantly, we show that our estimator remains root-n consistent and asymptotically normal so long as either the propensity score model or the outcome model is correctly specified. We provide valid confidence intervals in both cases and further extend these results to the case where the outcome model is a generalized linear model. In simulation studies, we find that the proposed methodology often estimates the average treatment effect more accurately than the existing methods. We also present an empirical application, in which we estimate the average causal effect of college attendance on adulthood political participation. Open-source software is available for implementing the proposed methodology.

研究动机与目标

  • 解决当混杂因素数量超过样本量时估计平均处理效应的挑战。
  • 开发一种即使倾向得分模型或结果模型误设仍保持一致性和有效性的方法。
  • 通过仅平衡最具预测力的协变量,提升高维观察性研究中估计的稳定性和效率。
  • 提供有效的置信区间,并将结果扩展至广义线性结果模型。
  • 避免依赖变量选择的一致性,转而专注于因果效应估计。

提出的方法

  • 使用用户指定的权重函数,通过惩罚M估计进行初始倾向得分估计。
  • 通过加权最小二乘法拟合结果模型,并使用独立的权重函数以增强稳健性。
  • 通过约束优化在结果预测性协变量子集上进行协变量平衡,从而改进倾向得分。
  • 使用平衡后的倾向得分进行逆概率加权,以估计平均处理效应。
  • 引入样本有界性以确保稳定性与有效推断。
  • 将方法扩展至广义线性模型以处理非线性结果,同时保持理论性质。

实验结果

研究问题

  • RQ1当倾向得分或结果模型任一误设时,高维倾向得分方法是否仍能保持根n一致性与渐近正态性?
  • RQ2在混杂因素数量超过样本量的高维设置中,协变量平衡如何有效应用?
  • RQ3在模型设定正确时,所提出的方法是否能达到半参数效率?
  • RQ4在模型误设情况下,是否可构建有效的置信区间而无需依赖变量选择的一致性?
  • RQ5在有限样本中,HD-CBPS与现有方法(如CBPS、AIPW和IPW)相比,在偏差、方差与覆盖区间方面表现如何?

主要发现

  • 当倾向得分或结果模型任一正确设定时,HD-CBPS实现根n一致性与渐近正态性,表现出双重稳健性。
  • 在倾向得分设定正确且结果模型为线性时,该方法保持半参数效率。
  • 模拟结果显示,HD-CBPS在估计ATE时表现优于现有方法,偏差更低、标准误更小,尤其在高维设置下优势明显。
  • 实证应用显示,大学入学对政治参与具有正向平均处理效应(ATE = 0.8293,SE = 0.1247),推断稳定。
  • CBPS在白人群体子样本(n=966)中无法收敛,而HD-CBPS保持稳定,且标准误小于AIPW-NR,结果可靠。
  • 该方法提供诚实的置信区间,对调优参数选择具有稳健性,敏感性分析已证实此点。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。