[Paper Review] The correlation-assisted missing data estimator
The paper proposes the Correlation-Assisted Missing data (CAM) estimator, a novel method that improves estimation accuracy by leveraging correlations between complete cases and incomplete observations. It reduces mean squared error compared to complete-case analysis in U-statistics, kernel density estimation, and nonparametric regression, with theoretical guarantees of lower variance and asymptotic normality under missing completely at random assumptions.
We introduce a novel approach to estimation problems in settings with missing data. Our proposal -- the Correlation-Assisted Missing data (CAM) estimator -- works by exploiting the relationship between the observations with missing features and those without missing features in order to obtain improved prediction accuracy. In particular, our theoretical results elucidate general conditions under which the proposed CAM estimator has lower mean squared error than the widely used complete-case approach in a range of estimation problems. We showcase in detail how the CAM estimator can be applied to $U$-Statistics to obtain an unbiased, asymptotically Gaussian estimator that has lower variance than the complete-case $U$-Statistic. Further, in nonparametric density estimation and regression problems, we construct our CAM estimator using kernel functions, and show it has lower asymptotic mean squared error than the corresponding complete-case kernel estimator. We also include practical demonstrations throughout the paper using simulated data and the Terneuzen birth cohort and Brandsma datasets available from CRAN.
Motivation & Objective
- To address the inefficiency of complete-case analysis in missing data settings, where valuable information from incomplete observations is discarded.
- To develop a general framework that improves estimation accuracy by exploiting correlations between complete cases and observations with missing features.
- To provide theoretical guarantees for the CAM estimator's lower mean squared error compared to complete-case estimators across multiple statistical models.
- To enable practical implementation through data-driven, fast, and fully automatic construction of the estimator in nonparametric and U-statistic settings.
Proposed method
- The CAM estimator constructs a linear adjustment to the complete-case estimator using a statistic correlated with it, derived from both complete and incomplete observations.
- It uses the correlation between the complete-case estimator and an auxiliary estimator based on incomplete data to reduce mean squared error.
- For U-statistics, the optimal adjustment term is also a U-statistic, enabling theoretical analysis and construction of an unbiased, asymptotically Gaussian CAM estimator.
- In nonparametric settings, the method applies kernel functions to construct the CAM estimator, with theoretical justification for asymptotic variance reduction.
- A data-driven version of the CAM estimator is proposed using empirical correlation estimates, ensuring practical usability without requiring knowledge of the optimal adjustment.
- Theoretical analysis includes bounds on bias and variance using concentration inequalities and asymptotic expansions under standard nonparametric regularity conditions.
Experimental results
Research questions
- RQ1Under what conditions does the CAM estimator achieve lower mean squared error than the complete-case estimator?
- RQ2Can the CAM estimator be constructed to be unbiased and asymptotically normal in U-statistic settings under missing completely at random mechanisms?
- RQ3How does the CAM estimator improve asymptotic mean squared error in kernel-based nonparametric density estimation and regression?
- RQ4What is the optimal adjustment term in the CAM estimator, and how can it be estimated from data in practice?
- RQ5Can the CAM estimator be applied effectively to real-world datasets with missing data, such as the Terneuzen birth cohort and Brandsma datasets?
Key findings
- The CAM estimator achieves lower mean squared error than the complete-case estimator under general conditions, without requiring the data to be missing completely at random.
- In U-statistic settings, the CAM estimator is unbiased and asymptotically Gaussian when data is missing completely at random, with asymptotic variance strictly smaller than that of the complete-case U-statistic.
- For kernel density estimation and local constant regression, the CAM estimator has lower asymptotic mean squared error than the complete-case kernel estimator, with the improvement quantified in leading-order terms.
- The optimal adjustment term in the CAM estimator corresponds to a U-statistic in the U-statistic setting, enabling theoretical analysis and construction of a consistent estimator.
- The data-driven CAM estimator achieves significant variance reduction in simulations and real data applications, including the Terneuzen birth cohort and Brandsma datasets.
- Theoretical bounds confirm that the CAM estimator’s bias and variance terms converge appropriately, with residual terms vanishing in probability under standard nonparametric regularity conditions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.