[论文解读] Propensity score regression for causal inference with treatment heterogeneity
本文提出了一种非参数倾向得分回归(PSR)方法,用于估计异质处理效应,该方法对极端倾向得分和高维混杂变量具有鲁棒性。通过结合两种非参数回归——首先对倾向得分和感兴趣的协变量进行回归,然后将结果仅对协变量进行回归——PSR实现了具显式方差估计器的一致性、渐近正态估计,在模拟和真实世界中对流感疫苗接种与病假影响的分析中优于现有方法。
Understanding how treatment effects vary on individual characteristics is critical in the contexts of personalized medicine, personalized advertising and policy design. When the characteristics are of practical interest are only a subset of full covariate, non-parametric estimation is often desirable; but few methods are available due to the computational difficult. Existing non-parametric methods such as the inverse probability weighting methods have limitations that hinder their use in many practical settings where the values of propensity scores are close to 0 or 1. We propose the propensity score regression (PSR) that allows the non-parametric estimation of the heterogeneous treatment effects in a wide context. PSR includes two non-parametric regressions in turn, where it first regresses on the propensity scores together with the characteristics of interest, to obtain an intermediate estimate; and then, regress the intermediate estimates on the characteristics of interest only. By including propensity scores as regressors in the non-parametric manner, PSR is capable of substantially easing the computational difficulty while remain (locally) insensitive to any value of propensity scores. We present several appealing properties of PSR, including the consistency and asymptotical normality, and in particular the existence of an explicit variance estimator, from which the analytical behaviour of PSR and its precision can be assessed. Simulation studies indicate that PSR outperform existing methods in varying settings with extreme values of propensity scores. We apply our method to the national 2009 flu survey (NHFS) data to investigate the effects of seasonal influenza vaccination and having paid sick leave across different age groups.
研究动机与目标
- 为解决在个性化医疗和政策设计中,基于年龄等关键特征而变化的条件平均处理效应估计挑战。
- 克服当倾向得分为接近零或一时,逆概率加权和AIPW方法存在的估计不稳定性问题。
- 开发一种非参数方法,灵活建模处理效应异质性,而无需对结果或倾向得分施加参数假设。
- 通过利用倾向得分的平衡性质,减轻高维设置下的计算负担并提高鲁棒性。
- 提供具有理论保证的方法,包括一致性、渐近正态性以及显式方差估计器,以支持推断。
提出的方法
- PSR方法采用两阶段非参数回归:首先,将结果对倾向得分和感兴趣的协变量进行回归,以获得中间估计量。
- 其次,将该中间估计量仅对感兴趣的协变量进行回归,从而有效消除高维混杂变量的影响。
- 该方法使用基于核的非参数回归,并通过带宽选择来估计条件期望,确保在建模异质效应时具有灵活性。
- 它利用倾向得分的平衡性质,即在每个得分水平上,处理组与对照组的协变量分布相等,从而控制来自高维$X^{-l}$的混杂影响。
- 该方法对极端倾向得分具有鲁棒性,因为它使用得分的连续有界函数,而非逆权重。
- 推导出显式方差估计器,从而实现分析精度评估,并支持对估计处理效应的有效推断。
实验结果
研究问题
- RQ1如何在对极端倾向得分保持鲁棒性的同时,非参数地估计在年龄等关键特征上变化的处理效应?
- RQ2使用倾向得分为回归变量的两阶段非参数回归方法,在有限样本和渐近性质方面如何?
- RQ3当倾向得分为接近零或一时,所提出的PSR方法与逆概率加权和AIPW方法在偏差、方差和鲁棒性方面有何比较?
- RQ4PSR方法能否在不遭受维度灾难影响的情况下,有效处理高维混杂变量?
- RQ5使用倾向得分的连续函数而非逆权重,对估计稳定性有何影响?
主要发现
- 在正则条件下,PSR方法实现了具显式方差估计器的一致性和渐近正态性,支持有效推断。
- 在模拟研究中,当倾向得分为接近零或一时,PSR优于现有方法,特别是逆概率加权和AIPW方法,表现出更低的偏差和均方误差。
- 即使在高维混杂变量下,该方法仍保持良好性能,通过利用倾向得分的平衡性质避免了维度灾难。
- 在NHFS应用中,PSR显示季节性流感疫苗接种会增加65岁以上人群的医生就诊次数,而有带薪病假者在33岁以下或41至64岁之间的人群中会减少医生就诊次数,33至40岁人群无显著影响。
- PSR方法对极端倾向得分具有鲁棒性,其在倾向得分范围从0.001到0.984的模拟中表现出稳定性能。
- 理论扩展表明,该方法在离散$X^l$和半参数模型(如单 index 模型)下依然有效,扩大了其适用范围。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。