Skip to main content
QUICK REVIEW

[Paper Review] A New Central Limit Theorem for the Augmented IPW Estimator: Variance Inflation, Cross-Fit Covariance and Beyond

Kuanhao Jiang, Rajarshi Mukherjee|arXiv (Cornell University)|May 20, 2022
Advanced Causal Inference Techniques4 citations
TL;DR

This paper establishes a new central limit theorem for the cross-fitted augmented inverse probability weighting (AIPW) estimator in high-dimensional settings without requiring sparsity assumptions. It reveals two key phenomena: substantial variance inflation dependent on the signal-to-noise ratio and non-negligible asymptotic covariance between pre-cross-fit estimators, both of which are absent in classical asymptotic theory. The proof leverages a novel combination of approximate message passing, deterministic equivalents, and leave-one-out analysis.

ABSTRACT

Estimation of the average treatment effect (ATE) is a central problem in causal inference. In recent times, inference for the ATE in the presence of high-dimensional covariates has been extensively studied. Among the diverse approaches that have been proposed, augmented inverse probability weighting (AIPW) with cross-fitting has emerged a popular choice in practice. In this work, we study this cross-fit AIPW estimator under well-specified outcome regression and propensity score models in a high-dimensional regime where the number of features and samples are both large and comparable. Under assumptions on the covariate distribution, we establish a new central limit theorem for the suitably scaled cross-fit AIPW that applies without any sparsity assumptions on the underlying high-dimensional parameters. Our CLT uncovers two crucial phenomena among others: (i) the AIPW exhibits a substantial variance inflation that can be precisely quantified in terms of the signal-to-noise ratio and other problem parameters, (ii) the asymptotic covariance between the pre-cross-fit estimators is non-negligible even on the root-n scale. These findings are strikingly different from their classical counterparts. On the technical front, our work utilizes a novel interplay between three distinct tools--approximate message passing theory, the theory of deterministic equivalents, and the leave-one-out approach. We believe our proof techniques should be useful for analyzing other two-stage estimators in this high-dimensional regime. Finally, we complement our theoretical results with simulations that demonstrate both the finite sample efficacy of our CLT and its robustness to our assumptions.

Motivation & Objective

  • To establish a new central limit theorem for the cross-fitted AIPW estimator in high-dimensional settings where the number of features and samples are large and comparable.
  • To characterize the asymptotic behavior of the AIPW estimator under well-specified outcome regression and propensity score models without assuming sparsity in high-dimensional parameters.
  • To uncover and quantify two novel phenomena: variance inflation and non-negligible asymptotic covariance between pre-cross-fit estimators.
  • To develop a proof framework combining approximate message passing, deterministic equivalents, and leave-one-out techniques for analyzing two-stage estimators in high-dimensional regimes.

Proposed method

  • Derives a new central limit theorem for the cross-fitted AIPW estimator under high-dimensional asymptotics with i.i.d. covariates having light-tailed distributions.
  • Utilizes approximate message passing theory to analyze the high-dimensional behavior of the ridge-regularized regression estimators used in the AIPW construction.
  • Applies the theory of deterministic equivalents to characterize the limiting behavior of quadratic forms involving high-dimensional regression coefficients.
  • Employs a leave-one-out approach to decouple dependencies and control estimation error in the two-stage AIPW framework.
  • Establishes asymptotic normality of the AIPW estimator under general moment conditions, without requiring sparsity or structural assumptions on the high-dimensional parameters.
  • Derives explicit expressions for the asymptotic variance and covariance structure, showing dependence on the signal-to-noise ratio and problem-specific parameters.

Experimental results

Research questions

  • RQ1How does the asymptotic distribution of the cross-fitted AIPW estimator behave in high-dimensional settings without sparsity assumptions?
  • RQ2What is the impact of high-dimensional estimation on the variance of the AIPW estimator, and can it be quantified?
  • RQ3Why do pre-cross-fit estimators in AIPW exhibit non-negligible asymptotic covariance on the √n scale, contrary to classical theory?
  • RQ4Can a new central limit theorem be established for AIPW that captures these high-dimensional phenomena?
  • RQ5What novel technical tools are required to analyze two-stage estimators like AIPW in high-dimensional, non-sparse regimes?

Key findings

  • The AIPW estimator exhibits significant variance inflation that is precisely quantifiable in terms of the signal-to-noise ratio and other problem parameters, deviating from classical behavior.
  • The asymptotic covariance between pre-cross-fit estimators is non-negligible on the √n scale, a phenomenon not captured by classical asymptotic theory.
  • The limiting distribution of the AIPW estimator is non-normal in the conventional sense and depends on the joint distribution of the nuisance parameter estimates.
  • The asymptotic variance of the AIPW estimator is inflated relative to the classical case, with the inflation factor depending on the ridge regularization parameter and the covariance structure of the design matrix.
  • The theoretical framework successfully captures the finite-sample behavior observed in simulations, demonstrating robustness to model misspecification and distributional assumptions.
  • The proof technique, combining approximate message passing, deterministic equivalents, and leave-one-out analysis, is generalizable to other two-stage estimators in high-dimensional settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.