Skip to main content
QUICK REVIEW

[Paper Review] Adaptive debiased machine learning using data-driven model selection techniques

Lars van der Laan, Marco Carone|arXiv (Cornell University)|Jul 24, 2023
Statistical Methods and InferenceMathematics3 citations
TL;DR

This paper proposes Adaptive Debiased Machine Learning (ADML), a nonparametric framework that combines data-driven model selection with debiased machine learning to produce asymptotically linear, superefficient estimators for pathwise differentiable functionals. By learning an oracle submodel from data, ADML achieves regular, locally uniformly valid inference for a projection-based oracle parameter that matches the target parameter when the true distribution lies within the learned submodel, ensuring no asymptotic loss from data-driven model selection compared to knowing the oracle model in advance.

ABSTRACT

Debiased machine learning estimators for smooth functionals in nonparametric models can exhibit substantial variability and instability, often leading practitioners to instead rely on parametric or semiparametric working models. Such models, however, may be misspecified and can therefore introduce bias. We study how data-driven model selection can be combined with debiased machine learning to construct estimators that adapt to structure in the data-generating distribution. To this end, we propose Adaptive Debiased Machine Learning (ADML), a nonparametric framework for constructing superefficient estimators of pathwise differentiable parameters. The framework unifies a broad class of previously proposed adaptive estimators, including methods based on variable selection, learned feature representations, and collaborative targeted learning. It requires only high-level conditions and approximate validity of the selection procedure, which are implied by lower-level conditions already assumed in important settings, including sieve-based selection, sparsity-based methods such as the Lasso, and data-adaptive feature representations. We show that ADML estimators yield regular and efficient root-\(n\) inference for an oracle projection parameter induced by a data-adaptive oracle submodel. This oracle parameter coincides with the target parameter at the true distribution but typically has a smaller efficiency bound, thereby yielding superefficiency for the target parameter. As a practical illustration, we introduce a broad class of automatic ADML estimators for continuous linear functionals of the outcome regression, in which model selection is performed directly on the regression itself. Motivated by overlap challenges in causal inference, we develop new superefficient plug-in estimators for the average treatment effect based on calibration in semiparametric regression models.

Motivation & Objective

  • To address the limitation of standard debiased machine learning estimators, which are not adaptive to unknown structural features like sparsity or smoothness in the true data-generating distribution.
  • To develop a framework that avoids parametric or semiparametric model assumptions while still achieving efficiency when such structure exists.
  • To ensure that data-driven model selection does not incur an asymptotic cost in bias or variance, even under local alternatives.
  • To provide theoretically grounded, regular inference for adaptive estimators that are superefficient under the oracle submodel.
  • To demonstrate practical applicability through ADML estimators for average treatment effect in adaptive partially linear models.

Proposed method

  • ADML uses data-driven model selection to learn a working model approximating an unknown oracle submodel of a prespecified statistical model.
  • It projects the true data-generating distribution onto the learned working submodel to define a data-adaptive estimand.
  • Debiased machine learning techniques are then applied to construct estimators for this projection-based estimand.
  • The framework ensures asymptotic linearity and local uniform validity of inference for the oracle parameter, which coincides with the target parameter under the oracle submodel.
  • Sample-splitting and cross-fitting techniques are recommended to reduce finite-sample variance inflation from model selection.
  • Bootstrap or subsampling methods may be used to improve variance estimation and confidence interval coverage in finite samples.

Experimental results

Research questions

  • RQ1Can data-driven model selection in debiased machine learning be performed without incurring an asymptotic cost in bias or variance?
  • RQ2Does the resulting estimator remain regular and provide valid inference under local asymptotic perturbations?
  • RQ3Is there a local asymptotic equivalence between using the oracle submodel (known in advance) and learning it from data?
  • RQ4How does ADML compare to fixed parametric or semiparametric estimators in terms of bias, variance, and coverage under model misspecification?
  • RQ5Can ADML achieve superefficiency while maintaining robustness to model misspecification?

Key findings

  • ADML estimators are asymptotically linear and provide locally uniformly valid inference for a projection-based oracle parameter that matches the target parameter when the true distribution lies within the learned oracle submodel.
  • There is no asymptotic loss in performing data-driven model selection compared to knowing the oracle submodel in advance, even under least-favorable local perturbations.
  • The prespecified semiparametric estimator (e.g., constant CATE) and the partially linear ADMLE were asymptotically equivalent under the least-favorable local alternative, supporting the no-loss result.
  • Confidence intervals based on the AIPW estimator achieved 95% coverage under local perturbations but at the cost of significantly increased variance and worse mean squared error in moderate overlap settings.
  • The plug-in HAL-ADMLE showed higher asymptotic bias than the partially linear ADMLE, consistent with its regularity under a smaller oracle submodel.
  • In extreme overlap settings (e.g., $c_0 \approx 10^{-6}$), the AIPW estimator became biased and highly variable, likely due to non-identifiability of the ATE in the nonparametric model.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.