Skip to main content
QUICK REVIEW

[论文解读] Robust Estimation of Propensity Score Weights via Subclassification

Linbo Wang, Zhang, Yuexia|arXiv (Cornell University)|Feb 20, 2016
Advanced Causal Inference Techniques参考文献 35被引用 5
一句话总结

本文提出全分类权重——通过基于排序倾向得分估计值对单位进行子分类而导出——以在观察性研究中构建稳健、模型无关的加权估计量,用于因果推断。通过用子组内经验均值替代参数化倾向得分加权,该方法在模型误设情况下实现了更优的稳定性与协变量平衡,同时确保了不同加权估计量的一致性与收敛性。

ABSTRACT

Weighting estimators based on propensity scores are widely used for causal estimation in a variety of contexts, such as observational studies, marginal structural models and interference. They enjoy appealing theoretical properties such as consistency and possible efficiency under correct model specification. However, this theoretical appeal may be diminished in practice by sensitivity to misspecification of the propensity score model. To improve on this, we borrow an idea from an alternative approach to causal effect estimation in observational studies, namely subclassification estimators. It is well known that compared to weighting estimators, subclassification methods are usually more robust to model misspecification. In this paper, we first discuss an intrinsic connection between the seemingly unrelated weighting and subclassification estimators, and then use this connection to construct robust propensity score weights via subclassification. We illustrate this idea by proposing so-called full-classification weights and accompanying estimators for causal effect estimation in observational studies. Our novel estimators are both consistent and robust to model misspecification, thereby combining the strengths of traditional weighting and subclassification estimators for causal effect estimation from observational studies. Numerical studies show that the proposed estimators perform favorably compared to existing methods.

研究动机与目标

  • 解决传统逆概率加权估计量在观察性研究中对倾向得分模型误设的敏感性问题。
  • 开发一种加权方法,使其在设计阶段不依赖结果数据,同时保持一致性和稳健性。
  • 结合加权方法的理论一致性与分层方法的实证稳健性。
  • 在倾向得分模型误设时,实现优于参数化加权的更好平衡性与权重稳定性。
  • 提供一个统一框架,使不同加权估计量(如Horvitz-Thompson、比率估计、双重稳健)在新权重下产生一致结果。

提出的方法

  • 基于参数模型估计的倾向得分的分位数,仅使用排序信息将单位划分为子组。
  • 在每个子组内计算经验均值作为权重,以非参数、基于排序的方法替代直接的逆概率加权。
  • 构建对初始倾向得分模型函数形式不敏感的全分类(FS)权重,聚焦于子组层面的平衡性。
  • 确保权重独立于结果数据,保持客观性,并可在多种因果推断场景中应用。
  • 使用基于自助法的推断方法,计算在所提加权方案下因果效应估计的置信区间。
  • 证明Horvitz-Thompson、比率估计和双重稳健估计量在全分类权重下收敛至相同估计值,从而增强稳健性。

实验结果

研究问题

  • RQ1对倾向得分估计值进行分层是否能提升基于加权的因果推断在模型误设情况下的稳健性?
  • RQ2全分类权重在权重稳定性与协变量平衡方面,相较于参数化、截断和协变量平衡加权方法的表现如何?
  • RQ3在全分类权重下,不同加权估计量(如HT、比率、DR)在因果效应估计上的一致性程度如何?
  • RQ4当初始倾向得分模型误设时,所提方法是否仍能保持平均处理效应估计的一致性?
  • RQ5全分类方法能否推广至其他场景,如边际结构模型或MAR下的缺失数据?

主要发现

  • 在模型误设情况下,全分类权重在协变量平衡性和权重稳定性方面显著优于参数化、截断及其他稳健加权方法。
  • 在全分类权重下,Horvitz-Thompson、比率估计和双重稳健估计量产生的因果效应估计几乎完全一致,表明其收敛性与稳健性。
  • 在NHANES数据分析中,所有稳健加权方法(包括全分类)均显示学校餐食参与对BMI的影响可忽略,而参数化模型则产生高度可变的估计结果。
  • 原始估计量(未调整)显示正向效应(1.04),但全分类及其他稳健方法的估计值接近零(如-0.20,95%置信区间:-0.75, 0.36),表明偏差显著降低。
  • 全分类方法实现了0.12的平衡偏差,与表现最佳的方法(如CAL-ET的0.00)相当,显示出强大的实证平衡性。
  • 该方法有效缓解了模型误设的影响:尽管不同参数化模型导致Horvitz-Thompson估计量结果迥异,但在全分类权重下,各模型均产生一致的估计结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。