Skip to main content
QUICK REVIEW

[Paper Review] Data-adaptive doubly robust instrumental variable methods for treatment effect heterogeneity

Karla Díaz-Ordaz, Rhian Daniel|arXiv (Cornell University)|Feb 8, 2018
Advanced Causal Inference Techniques48 references3 citations
TL;DR

This paper proposes data-adaptive doubly robust instrumental variable estimators—specifically a locally efficient g-estimator and a targeted maximum likelihood estimator (TMLE)—for estimating treatment effect heterogeneity in randomized trials with non-adherence. By integrating machine learning via Super Learner to estimate nuisance parameters, the methods achieve lower bias and improved coverage than traditional two-stage least squares (TSLS) or parametric doubly robust estimators, especially under model misspecification.

ABSTRACT

We consider the estimation of the average treatment effect in the treated as a function of baseline covariates, where there is a valid (conditional) instrument. We describe two doubly robust (DR) estimators: a locally efficient g-estimator, and a targeted minimum loss-based estimator (TMLE). These two DR estimators can be viewed as generalisations of the two-stage least squares (TSLS) method to semi-parametric models that make weaker assumptions. We exploit recent theoretical results that extend to the g-estimator the use of data-adaptive fits for the nuisance parameters. A simulation study is used to compare standard TSLS with the two DR estimators' finite-sample performance, (1) when fitted using parametric nuisance models, and (2) using data-adaptive nuisance fits, obtained from the Super Learner, an ensemble machine learning method. Data-adaptive DR estimators have lower bias and improved coverage, when compared to incorrectly specified parametric DR estimators and TSLS. When the parametric model for the treatment effect curve is correctly specified, the g-estimator outperforms all others, but when this model is misspecified, TMLE performs best, while TSLS can result in large biases and zero coverage. Finally, we illustrate the methods by reanalysing the COPERS (COping with persistent Pain, Effectiveness Research in Self-management) trial to make inference about the causal effect of treatment actually received, and the extent to which this is modified by depression at baseline.

Motivation & Objective

  • To address causal inference in randomized trials with non-adherence, where treatment received may differ from treatment assigned.
  • To estimate the average treatment effect in the treated as a function of baseline covariates, accounting for treatment effect heterogeneity.
  • To improve estimation efficiency and robustness by leveraging data-adaptive methods for nuisance parameter estimation in instrumental variable models.
  • To evaluate the finite-sample performance of doubly robust estimators when fitted with parametric versus machine learning-based nuisance models.
  • To apply and illustrate the methods in a real-world trial (COPERS) to assess the causal effect of pain self-management interventions modified by baseline depression.

Proposed method

  • Proposes two doubly robust (DR) estimators: a locally efficient g-estimator and a targeted maximum likelihood estimator (TMLE) for average treatment effect in the treated with a valid instrument.
  • Extends the doubly robust property to settings with treatment effect heterogeneity by modeling the outcome and exposure conditional on baseline covariates.
  • Uses data-adaptive estimation via the Super Learner ensemble machine learning method to estimate nuisance parameters (e.g., outcome regression, propensity scores, instrument models).
  • Employs orthogonality and influence function-based estimation to ensure asymptotic normality and valid inference even when machine learning estimators converge slowly.
  • Applies targeted minimum loss-based estimation (TMLE) with cross-validated loss minimization to improve efficiency and bias reduction in finite samples.
  • Derives regularity conditions under which the data-adaptive g-estimator is asymptotically linear and root-n consistent, relying on convergence rates of nuisance estimators.

Experimental results

Research questions

  • RQ1How do data-adaptive doubly robust estimators compare to standard two-stage least squares (TSLS) in terms of bias and coverage under model misspecification?
  • RQ2What is the finite-sample performance of doubly robust estimators when nuisance parameters are estimated using parametric models versus machine learning (e.g., Super Learner)?
  • RQ3Does the choice of estimator (g-estimator vs. TMLE) matter when the treatment effect model is correctly specified versus misspecified?
  • RQ4Can data-adaptive DR estimators maintain low bias and good coverage when both the outcome and exposure models are misspecified?
  • RQ5To what extent does baseline depression modify the causal effect of treatment received in the COPERS trial?

Key findings

  • Data-adaptive doubly robust estimators using Super Learner for nuisance parameters showed significantly lower bias and improved coverage compared to parametric doubly robust estimators and TSLS.
  • When the treatment effect model was correctly specified, the g-estimator outperformed all other methods in terms of bias and mean squared error.
  • When the treatment effect model was misspecified, the TMLE-based estimator performed best, while TSLS produced large biases and zero coverage in some scenarios.
  • In a high-sample-size scenario ($n=10,000$) with both outcome and exposure models misspecified, TSLS had a coverage rate of 0.000, indicating severe inference failure.
  • The TMLE-based estimator achieved 100% coverage in the same high-n, misspecified scenario, demonstrating robustness under model misspecification.
  • In the COPERS trial reanalysis, the causal effect of treatment received on pain-related disability was found to be modified by baseline depression, with data-adaptive DR methods providing reliable inference.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.