Skip to main content
QUICK REVIEW

[Paper Review] Nonparametric Pattern-Mixture Models for Inference with Missing Data

Yen‐Chi Chen, Mauricio Sadinle|arXiv (Cornell University)|Apr 24, 2019
Bayesian Methods and Mixture Models52 references4 citations
TL;DR

This paper proposes a nonparametric pattern-mixture model for inference with missing data, using a surrogate empirical distribution with kernel smoothing and Monte Carlo sampling under donor-based identifying restrictions. It enables valid point estimation and bootstrap confidence intervals for functionals like ATEs, even under nonmonotone missingness, with theoretical guarantees and robustness to misspecified assumptions.

ABSTRACT

Pattern-mixture models provide a transparent approach for handling missing data, where the full-data distribution is factorized in a way that explicitly shows the parts that can be estimated from observed data alone, and the parts that require identifying restrictions. We introduce a nonparametric estimator of the full-data distribution based on the pattern-mixture model factorization. Our approach uses the empirical observed-data distribution and augments it with a nonparametric estimator of the missing-data distributions under a given identifying restriction. Our results apply to a large class of donor-based identifying restrictions that encompasses commonly used ones and can handle both monotone and nonmonotone missingness. We propose a Monte Carlo procedure to derive point estimates of functionals of interest, and the bootstrap to construct confidence intervals.

Motivation & Objective

  • To develop a nonparametric framework for pattern-mixture models that handles both monotone and nonmonotone missingness.
  • To unify and generalize donor-based identifying restrictions for use in nonparametric settings.
  • To provide a computationally feasible method for estimating functionals of interest under missing data, such as average treatment effects.
  • To establish theoretical validity of Monte Carlo estimation and bootstrap confidence intervals under nonparametric estimation.
  • To enable sensitivity analysis by allowing multiple identifying restrictions to be evaluated under a single consistent framework.

Proposed method

  • The method uses the empirical distribution of observed data as a base and nonparametrically estimates missing-data distributions using conditional kernel-density estimators.
  • It introduces a surrogate observed-data distribution based on kernel smoothing to enable estimation under donor-based identifying restrictions.
  • A Monte Carlo procedure is used to simulate from the estimated full-data distribution, enabling numerical computation of functionals like ATEs.
  • The bootstrap is applied to construct valid 95% confidence intervals for functionals, with theoretical justification provided via Efron’s bootstrap diagram.
  • The approach is designed to be robust: inferences depend only on observed data and a specified identifying restriction, even if the restriction is misspecified.
  • Convergence rates and asymptotic normality are derived under regularity conditions, supporting inferential validity.

Experimental results

Research questions

  • RQ1How can nonparametric inference be conducted under pattern-mixture models when missingness is nonmonotone?
  • RQ2Can a unified framework for donor-based identification be extended to handle nonmonotone missing data?
  • RQ3What is the theoretical validity of Monte Carlo approximation for functionals under nonparametric pattern-mixture models?
  • RQ4How can bootstrap-based confidence intervals be constructed and justified under this nonparametric framework?
  • RQ5To what extent are inferences robust to misspecification of identifying restrictions?

Key findings

  • The proposed Monte Carlo method for estimating functionals is theoretically valid, with convergence rates and asymptotic normality established under regularity conditions.
  • The bootstrap procedure for constructing confidence intervals is valid, as demonstrated through theoretical proof and Efron’s bootstrap diagram.
  • In a clinical trial on schizophrenia, all 95% bootstrap confidence intervals for the ATE comparing new vs. standard treatment at week 8 fell below zero, indicating strong evidence for superiority of the new treatment across all missing-data assumptions.
  • The method successfully handles nonmonotone missingness, extending the applicability of pattern-mixture models beyond longitudinal monotone settings.
  • The approach allows for sensitivity analysis by evaluating the same functional under multiple identifying restrictions (e.g., AC, 3NC, NC), with consistent conclusions across assumptions.
  • Theoretical results show that inferences depend only on observed data and the specified identifying restriction, ensuring transparency and interpretability.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.