Skip to main content
QUICK REVIEW

[Paper Review] Minimax Semiparametric Learning With Approximate Sparsity

Jelena Bradić, Victor Chernozhukov|arXiv (Cornell University)|Dec 27, 2019
Statistical Methods and Inference4 citations
TL;DR

This paper establishes minimax optimal conditions for root-n consistent estimation of high-dimensional, approximately sparse regression functionals—such as regression coefficients, average derivatives, and average treatment effects—using debiased machine learning with Lasso-based estimation and cross-fitting. The key contribution is showing that root-n consistency is achievable under minimal approximate sparsity ($\xi_1 > 1/2$) or rate double robustness ($\xi_1\xi_2 > 1/4$), extending prior results to broader conditions with unknown best regressor sets.

ABSTRACT

Estimating linear, mean-square continuous functionals is a pivotal challenge in statistics. In high-dimensional contexts, this estimation is often performed under the assumption of exact model sparsity, meaning that only a small number of parameters are precisely non-zero. This excludes models where linear formulations only approximate the underlying data distribution, such as nonparametric regression methods that use basis expansion such as splines, kernel methods or polynomial regressions. Many recent methods for root-$n$ estimation have been proposed, but the implications of exact model sparsity remain largely unexplored. In particular, minimax optimality for models that are not exactly sparse has not yet been developed. This paper formalizes the concept of approximate sparsity through classical semi-parametric theory. We derive minimax rates under this formulation for a regression slope and an average derivative, finding these bounds to be substantially larger than those in low-dimensional, semi-parametric settings. We identify several new phenomena. We discover new regimes where rate double robustness does not hold, yet root-$n$ estimation is still possible. In these settings, we propose an estimator that achieves minimax optimal rates. Our findings further reveal distinct optimality boundaries for ordered versus unordered nonparametric regression estimation.

Motivation & Objective

  • To establish necessary and sufficient conditions for root-n consistent estimation of linear, mean-square continuous functionals in high-dimensional, approximately sparse regression models.
  • To extend existing debiased machine learning estimators to achieve root-n consistency under weaker and more general conditions than previously known.
  • To clarify the role of approximate sparsity in determining feasibility of root-n estimation, particularly when the identity of the best regressors is unknown.
  • To develop estimators that maintain root-n consistency under minimal assumptions, including non-Gaussian errors and heteroskedasticity.
  • To demonstrate that the Riesz representer's sparsity rate ($\xi_2$) can alone ensure root-n consistency when $\xi_2 > 1/2$, even if the regression is dense.

Proposed method

  • Proposes a debiased machine learning estimator using Lasso for regression and bias correction, with special cross-fitting to ensure asymptotic normality.
  • Employs a two-stage estimation procedure: first estimate the regression function $\rho_0$ via Lasso, then estimate the Riesz representer $\alpha_0$ via Lasso on the influence function.
  • Uses cross-fitting with split samples to decouple estimation and estimation error, reducing bias in the final estimator.
  • Applies a truncation operator $\tau_n$ to control the $L_1$-norm of the estimated Riesz representer, ensuring stability.
  • Derives theoretical bounds on estimation error using sparse approximation rates $s^{-\xi}$ for both the regression and Riesz representer.
  • Establishes asymptotic normality and consistent variance estimation under minimal conditions, including $\xi_1 > 1/2$ or $\xi_1\xi_2 > 1/4$.

Experimental results

Research questions

  • RQ1What are the minimal conditions on the sparse approximation rates $\xi_1$ and $\xi_2$ that allow root-n consistent estimation of a regression slope or average derivative in high-dimensional models?
  • RQ2How does the feasibility of root-n estimation change when the identity of the best $s$ regressors is unknown versus known?
  • RQ3Can root-n consistency be achieved under rate double robustness ($\xi_1\xi_2 > 1/4$) without requiring $\xi_1 > 1/2$?
  • RQ4Is it possible to achieve root-n consistency when the regression is dense ($\xi_1 \leq 1/2$) but the Riesz representer is sparse ($\xi_2 > 1/2$)?
  • RQ5What is the role of cross-fitting and truncation in ensuring root-n consistency and asymptotic normality in high-dimensional semiparametric models?

Key findings

  • The necessary and sufficient condition for root-n consistency of a regression slope or average derivative is $\max\{\xi_1, \xi_2\} > 1/2$, which is stronger than the known-identity case.
  • Root-n consistency is achievable under the minimal condition $\xi_1 > 1/2$, even when the regression is dense, provided the Riesz representer is sufficiently sparse.
  • The rate double robustness condition $\xi_1\xi_2 > 1/4$ is sufficient for root-n consistency and allows for broader applicability than prior conditions.
  • An estimator without cross-fitting achieves root-n consistency under $\xi_1 > 1/2$ and additional regularity conditions, including $\bar{\tau}_n \to \infty$ and $\|\rho_n - \rho_0\|_2 \to 0$, with consistent variance estimation.
  • The wedge between the known-identity and unknown-identity conditions in Figure 1 confirms that estimation is strictly more difficult when the best regressors are unknown.
  • The proposed estimators achieve asymptotic normality $\sqrt{n}(\hat{\theta} - \theta_0) \overset{d}{\longrightarrow} N(0, V)$ and consistent variance estimation under minimal sparsity assumptions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.