[论文解读] Explaining Practical Differences Between Treatment Effect Estimators with High Dimensional Asymptotics
本文解释了为何在经典大样本理论中渐近效率等价的常用因果效应估计量——G-computation、IPW、AIPW 和 TMLE——在实践中表现出不同的方差。通过采用高维渐近理论(即混杂变量数 d 随样本量 n 增长),作者表明机器学习方法估计的干扰参数会引入不消失的方差,揭示了在低偏差条件下,G-computation 和 TMLE 由于在高维情形下具有更优的高阶渐近性质,因此优于其他估计量。
We revisit the classical causal inference problem of estimating the average treatment effect in the presence of fully observed confounding variables using two-stage semiparametric methods. In existing theoretical studies of methods such as G-computation, inverse propensity weighting (IPW), and two common doubly robust estimators -- augmented IPW (AIPW) and targeted maximum likelihood estimation (TMLE) -- they are either bias-dominated, or have similar asymptotic statistical properties. However, when applied to real datasets, they often appear to have notably different variance. We compare these methods when using a machine learning (ML) model to estimate the nuisance parameters of the semiparametric model, and highlight some of the important differences. When the outcome model estimates have little bias, which is common among some key ML models, G-computation and the TMLE outperforms the other estimators in both bias and variance. We show that the differences can be explained using high-dimensional statistical theory, where the number of confounders $d$ is of the same order as the sample size $n$. To make this theoretical problem tractable, we posit a generalized linear model for the effect of the confounders on the treatment assignment and outcomes. Despite making parametric assumptions, this setting is a useful surrogate for some machine learning methods used to adjust for confounding in two-stage semiparametric methods. In particular, the estimation of the first stage adds variance that does not vanish, forcing us to confront terms in the asymptotic expansion that normally are brushed aside as finite sample defects. However, our model emphasizes differences in performance between these estimators beyond first-order asymptotics.
研究动机与目标
- 解释在现实应用中,常见因果效应估计量(G-computation、IPW、AIPW、TMLE)之间实际表现差异的原因。
- 研究为何这些估计量在有限样本中表现出不同的方差,尽管在经典理论中渐近效率相近。
- 构建一个高维渐近框架,其中 d/n → κ ∈ (0,1),使分析可处理,同时捕捉机器学习方法估计干扰参数的关键特征。
- 识别在使用机器学习模型进行混杂变量调整时,G-computation 和 TMLE 在偏差与方差方面优于 AIPW 和 IPW 的条件。
- 证明在经典理论中常被忽略的高阶渐近项,对理解估计量实际表现差异至关重要。
提出的方法
- 提出一个高维渐近模型,其中混杂变量数 d 随样本量 n 成比例增长(d/n → κ ∈ (0,1))。
- 假设结果和处理机制的参数模型:E[W|X] = h⁻¹(ηᵀX),E[Y(1)|X] = g⁻¹(β₁ᵀX),E[Y(0)|X] = g⁻¹(β₀ᵀX)。
- 为简化和便于分析,采用线性模型(恒等链接),同时保持干扰参数大小非退化(ηᵀX = Θₚ(1))。
- 通过全期望公式并条件于估计的干扰参数,分解余项以分析估计量的渐近方差。
- 推导出原始估计量与 AIPW 估计量之间差异的渐近方差的显式表达式,表明其依赖于干扰参数估计量的方差。
- 应用逆威沙特分布理论,计算逆 Gram 矩阵的期望,这是推导干扰参数估计量渐近方差的关键。
实验结果
研究问题
- RQ1为何在实践中,G-computation 和 TMLE 通常优于 AIPW 和 IPW,尽管在大样本效率上理论上等价?
- RQ2在高维设置下,基于机器学习的干扰参数估计如何影响因果效应估计量的有限样本方差?
- RQ3在经典理论中通常被忽略的高阶渐近项,在解释估计量之间实际差异中起什么作用?
- RQ4在何种高维渐近框架下,双 robust 估计量的渐近性质会偏离其经典大样本行为?
- RQ5当干扰参数通过机器学习模型估计时,样本与协变量维度之比 d/n 如何影响不同因果效应估计量的相对表现?
主要发现
- 在 d/n → κ ∈ (0,1) 的高维渐近框架下,干扰参数估计的方差不会消失,引入了非渐近效应,从而解释了实际表现差异。
- 当结果模型估计偏差较低时(某些机器学习模型中常见),G-computation 和 TMLE 的方差低于 AIPW 和 IPW。
- 原始估计量与 AIPW 估计量之间差异的渐近方差,由干扰参数估计量的方差与影响函数方差乘积的迹决定。
- 逆 Gram 矩阵 ∑XᵢXᵢᵀ⁻¹ 的期望被推导为 Σ⁻¹/(N₁w − d − 1),量化了在高维抽样下干扰估计量的偏差。
- 高斯分布的对称性使得能够推导出对称随机向量上函数乘积期望的精确表达式,从而实现对余项的分析。
- 分析表明,AIPW 的表现对干扰参数估计量的方差敏感,而 G-computation 和 TMLE 在低偏差机器学习估计下更具鲁棒性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。