[论文解读] Debiased Inverse Propensity Score Weighting for Estimation of Average Treatment Effects with High-Dimensional Confounders
本文提出去偏逆倾向得分加权(DIPW)方法,用于在高维观察性研究中估计平均处理效应,其中倾向得分服从稀疏逻辑模型,而结果回归函数可为任意复杂形式。在温和条件下,该方法实现 $√n$-一致性与半参数效率,即使结果模型设定错误或为非参数模型,也能通过置信区间进行有效推断。
We consider estimation of average treatment effects given observational data with high-dimensional pretreatment variables. Existing methods for this problem typically assume some form of sparsity for the regression functions. In this work, we introduce a debiased inverse propensity score weighting (DIPW) scheme for average treatment effect estimation that delivers $\sqrt{n}$-consistent estimates when the propensity score follows a sparse logistic regression model; the outcome regression functions are permitted to be arbitrarily complex. We further demonstrate how confidence intervals centred on our estimates may be constructed. Our theoretical results quantify the price to pay for permitting the regression functions to be unestimable, which shows up as an inflation of the variance of the estimator compared to the semiparametric efficient variance by a constant factor, under mild conditions. We also show that when outcome regressions can be estimated faster than a slow $1/\sqrt{ \log n}$ rate, our estimator achieves semiparametric efficiency. As our results accommodate arbitrary outcome regression functions, averages of transformed responses under each treatment may also be estimated at the $\sqrt{n}$ rate. Thus, for example, the variances of the potential outcomes may be estimated. We discuss extensions to estimating linear projections of the heterogeneous treatment effect function and explain how propensity score models with more general link functions may be handled within our framework. An R package exttt{dipw} implementing our methodology is available on CRAN.
研究动机与目标
- 解决在高维观察性研究中,当结果回归函数复杂或设定错误时,估计平均处理效应的挑战。
- 开发一种方法,即使在结果回归函数不稀疏的情况下,也能保持 $√n$-一致性。
- 在对结果回归模型假设最少的条件下,构建处理效应的有效置信区间。
- 将框架扩展至估计潜在结果的泛函,如方差和异质处理效应的线性投影。
- 证明当结果回归可比 $1/\sqrt{\log n}$ 更快估计时,DIPW 实现半参数效率。
提出的方法
- 提出一种去偏逆倾向得分加权(DIPW)方案,通过基于倾向得分模型的双重稳健校正,修正逆倾向得分加权中的偏差。
- 对倾向得分 $\pi(x) = \mathbb{P}(T=1|X=x)$ 使用稀疏高维逻辑回归模型,其中 $s_\pi = o(\sqrt{n}/\log p)$。
- 采用非参数或灵活估计器 $\tilde{\mu}(x)$ 估计结果回归函数 $\mathbb{E}[Y|X=x]$,可通过 Lasso、随机森林或其他方法实现。
- 应用源自倾向得分估计器影响函数的去偏校正项,以消除加权估计量中的偏差。
- 在温和正则条件下,利用估计量的渐近正态性,以 DIPW 估计量为中心构建置信区间。
- 将框架扩展至估计潜在结果的泛函,如 $\mathbb{E}[h(Y(t))|X=x]$(对任意可测函数 $h$),包括方差和分位数。
实验结果
研究问题
- RQ1当仅假设倾向得分稀疏时,能否在结果回归函数任意复杂的情况下实现 $\sqrt{n}$-一致性估计?
- RQ2DIPW 估计量的渐近方差是多少?与半参数效率界相比如何?
- RQ3在何种条件下,DIPW 估计量可实现半参数效率?
- RQ4当结果模型设定错误或难以估计时,DIPW 相较于 AIPW 和 TMLE 表现如何?
- RQ5DIPW 框架能否扩展至估计均值以外的潜在结果泛函,如方差或分位数?
主要发现
- 当倾向得分服从稀疏逻辑模型时,DIPW 估计量即使在结果回归函数任意复杂的情况下,仍可实现 $\sqrt{n}$-一致性。
- 在温和正则条件下,DIPW 估计量的渐近方差相比半参数效率界被放大一个常数因子。
- 当结果回归函数可比 $1/\sqrt{\log n}$ 更快估计时,DIPW 估计量实现半参数效率。
- 实证结果表明,在结果模型密集或设定错误时,DIPW 在重叠性较差的场景中优于 AIPW 和 TMLE。
- 在具有复杂结果函数的高维挑战性场景中,基于随机森林的 DIPW 变体表现与基于 Lasso 的 DIPW 相当或更优。
- 该方法可实现潜在结果方差等泛函的 $\sqrt{n}$ 速率估计,如 $\mathrm{Var}(Y(1))$。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。