Skip to main content
QUICK REVIEW

[Paper Review] Nonparametric inference about mean functionals of nonignorable nonresponse data without identifying the joint distribution

Wei Li, Wang Miao|arXiv (Cornell University)|Oct 12, 2021
Statistical Methods and Inference4 citations
TL;DR

This paper proposes a nonparametric method for estimating mean functionals in the presence of nonignorable missing data using a shadow variable, avoiding identification of the full joint distribution. It establishes identification and asymptotic normality via a sieve-based estimator that solves a representer equation, achieving semiparametric efficiency under regularity conditions.

ABSTRACT

We consider identification and inference about mean functionals of observed covariates and an outcome variable subject to nonignorable missingness. By leveraging a shadow variable, we establish a necessary and sufficient condition for identification of the mean functional even if the full data distribution is not identified. We further characterize a necessary condition for $\sqrt{n}$-estimability of the mean functional. This condition naturally strengthens the identifying condition, and it requires the existence of a function as a solution to a representer equation that connects the shadow variable to the mean functional. Solutions to the representer equation may not be unique, which presents substantial challenges for nonparametric estimation and standard theories for nonparametric sieve estimators are not applicable here. We construct a consistent estimator for the solution set and then adapt the theory of extremum estimators to find from the estimated set a consistent estimator for an appropriately chosen solution. The estimator is asymptotically normal, locally efficient and attains the semiparametric efficiency bound under certain regularity conditions. We illustrate the proposed approach via simulations and a real data application on home pricing.

Motivation & Objective

  • Address the challenge of estimating mean functionals when outcome data are missing not at random (MNAR), where standard methods fail due to selection bias.
  • Overcome identification difficulties in MNAR settings by leveraging a shadow variable—fully observed, associated with the outcome but independent of missingness given covariates and outcome.
  • Develop a model-free estimation framework that avoids parametric assumptions about the missingness mechanism or outcome regression.
  • Establish conditions under which the mean functional is $√n$-estimable and the proposed estimator is consistent and asymptotically normal.
  • Achieve semiparametric efficiency by constructing a consistent estimator for the solution set of a representer equation and selecting an appropriate solution.

Proposed method

  • Use a shadow variable to induce a necessary and sufficient condition for identification of the mean functional, even when the full data distribution is not identified.
  • Formulate a representer equation—a Fredholm integral equation of the first kind—that links the shadow variable to the mean functional, serving as the core identification device.
  • Construct a sieve-based estimator to consistently estimate the solution set of the representer equation, addressing non-uniqueness of solutions.
  • Adapt extremum estimation theory to select a consistent solution from the estimated solution set, enabling inference on the mean functional.
  • Establish asymptotic normality and semiparametric efficiency under regularity conditions, including completeness and integrability assumptions on conditional densities.
  • Use singular value decomposition of the conditional expectation operator to characterize existence of solutions to the representer equation.

Experimental results

Research questions

  • RQ1Under what conditions can the mean functional of an outcome with nonignorable missingness be nonparametrically identified without identifying the full joint distribution?
  • RQ2What is the necessary and sufficient condition for $√n$-estimability of the mean functional in the presence of nonignorable missingness?
  • RQ3How can one consistently estimate the solution set of a representer equation when solutions are not unique, and how can a consistent estimator for the mean functional be constructed from this set?
  • RQ4Can the proposed estimator achieve the semiparametric efficiency bound under a key submodel where the representer solution is unique?
  • RQ5What are the regularity conditions ensuring the existence of solutions to the representer equation, and how do they relate to the completeness of the conditional expectation operator?

Key findings

  • The mean functional is nonparametrically identifiable under a necessary and sufficient condition involving the shadow variable, even when the full data distribution is not identified.
  • A necessary condition for $√n$-estimability is the existence of a solution to the representer equation, which is a Fredholm integral equation of the first kind.
  • The proposed sieve-based estimator is consistent for the mean functional, even though the solution to the representer equation is not unique.
  • The estimator is asymptotically normal and attains the semiparametric efficiency bound under regularity conditions, particularly in a submodel where the representer solution is unique.
  • Existence of solutions to the representer equation is guaranteed under completeness conditions and integrability assumptions on the conditional densities, as formalized via singular value decomposition.
  • The method provides a robust alternative to standard measurement error models by avoiding strong identifying assumptions, relying instead on minimal, empirically plausible shadow variable assumptions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.