Skip to main content
QUICK REVIEW

[Paper Review] Explaining Practical Differences Between Treatment Effect Estimators with High Dimensional Asymptotics

Steve Yadlowsky|arXiv (Cornell University)|Mar 23, 2022
Advanced Causal Inference Techniques4 citations
TL;DR

This paper explains why commonly used treatment effect estimators—G-computation, IPW, AIPW, and TMLE—exhibit practical differences in variance despite having equivalent asymptotic efficiency in classical large-sample theory. Using high-dimensional asymptotics where the number of confounders d scales with sample size n, the authors show that machine learning-based nuisance parameter estimation introduces non-vanishing variance, revealing that G-computation and TMLE outperform others when bias is low, due to favorable higher-order asymptotic properties in high-dimensional regimes.

ABSTRACT

We revisit the classical causal inference problem of estimating the average treatment effect in the presence of fully observed confounding variables using two-stage semiparametric methods. In existing theoretical studies of methods such as G-computation, inverse propensity weighting (IPW), and two common doubly robust estimators -- augmented IPW (AIPW) and targeted maximum likelihood estimation (TMLE) -- they are either bias-dominated, or have similar asymptotic statistical properties. However, when applied to real datasets, they often appear to have notably different variance. We compare these methods when using a machine learning (ML) model to estimate the nuisance parameters of the semiparametric model, and highlight some of the important differences. When the outcome model estimates have little bias, which is common among some key ML models, G-computation and the TMLE outperforms the other estimators in both bias and variance. We show that the differences can be explained using high-dimensional statistical theory, where the number of confounders $d$ is of the same order as the sample size $n$. To make this theoretical problem tractable, we posit a generalized linear model for the effect of the confounders on the treatment assignment and outcomes. Despite making parametric assumptions, this setting is a useful surrogate for some machine learning methods used to adjust for confounding in two-stage semiparametric methods. In particular, the estimation of the first stage adds variance that does not vanish, forcing us to confront terms in the asymptotic expansion that normally are brushed aside as finite sample defects. However, our model emphasizes differences in performance between these estimators beyond first-order asymptotics.

Motivation & Objective

  • To explain the practical performance differences among common treatment effect estimators (G-computation, IPW, AIPW, TMLE) in real-world applications.
  • To investigate why these estimators show varying variance in finite samples despite similar asymptotic efficiency in classical theory.
  • To develop a high-dimensional asymptotic framework where d/n → κ ∈ (0,1), making the analysis tractable while capturing key features of machine learning-based nuisance estimation.
  • To identify conditions under which G-computation and TMLE outperform AIPW and IPW in terms of bias and variance when using ML models for confounder adjustment.
  • To demonstrate that higher-order asymptotic terms, often ignored in classical theory, are critical for understanding practical differences in estimator performance.

Proposed method

  • Proposes a high-dimensional asymptotic model where the number of confounders d grows proportionally with sample size n (d/n → κ ∈ (0,1)).
  • Assumes parametric models for the outcome and treatment mechanisms: E[W|X] = h⁻¹(ηᵀX), E[Y(1)|X] = g⁻¹(β₁ᵀX), E[Y(0)|X] = g⁻¹(β₀ᵀX).
  • Uses linear models (identity link) for simplicity and tractability, while maintaining non-degenerate nuisance parameter magnitudes (ηᵀX = Θₚ(1)).
  • Analyzes the asymptotic variance of estimators by decomposing the remainder term using the law of total variance and conditioning on estimated nuisance parameters.
  • Derives explicit expressions for the asymptotic variance of the difference between oracle and AIPW estimators, showing dependence on the variance of nuisance parameter estimates.
  • Applies inverse Wishart distribution theory to compute the expectation of the inverse Gram matrix, crucial for deriving the asymptotic variance of the nuisance estimators.

Experimental results

Research questions

  • RQ1Why do G-computation and TMLE often outperform AIPW and IPW in practice, despite theoretical equivalence in large-sample efficiency?
  • RQ2How do machine learning-based nuisance parameter estimates affect the finite-sample variance of treatment effect estimators in high-dimensional settings?
  • RQ3What role do higher-order asymptotic terms—typically ignored in classical theory—play in explaining practical differences between estimators?
  • RQ4In what high-dimensional regime do the asymptotic properties of doubly robust estimators diverge from their classical large-sample behavior?
  • RQ5How does the aspect ratio d/n influence the relative performance of different treatment effect estimators when nuisance parameters are estimated via ML models?

Key findings

  • In high-dimensional asymptotics with d/n → κ ∈ (0,1), the variance of nuisance parameter estimates does not vanish, introducing non-asymptotic effects that explain practical performance differences.
  • When the outcome model is estimated with low bias (common in some ML models), G-computation and TMLE achieve lower variance than AIPW and IPW.
  • The asymptotic variance of the difference between the oracle and AIPW estimators is driven by the trace of the product of the variance of the nuisance parameter estimators and the variance of the influence functions.
  • The expectation of the inverse Gram matrix ∑XᵢXᵢᵀ⁻¹ is derived as Σ⁻¹/(N₁w − d − 1), which quantifies the bias in the nuisance estimator under high-dimensional sampling.
  • The symmetry of the Gaussian distribution allows the derivation of an exact expression for the expectation of a product of functions over symmetric random vectors, enabling the analysis of the remainder term.
  • The analysis reveals that AIPW’s performance is sensitive to the variance of the nuisance parameter estimates, while G-computation and TMLE are more robust under low-bias ML estimation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.