Skip to main content
QUICK REVIEW

[Paper Review] Double Robust Semi-Supervised Inference for the Mean: Selection Bias under MAR Labeling with Decaying Overlap

Yuqian Zhang, Abhishek Chakrabortty|arXiv (Cornell University)|Apr 14, 2021
Advanced Causal Inference Techniques39 references4 citations
TL;DR

This paper proposes a double robust semi-supervised (DRSS) estimator for the mean under missing at random (MAR) labeling with decaying overlap, where labeled data are scarce and selection bias arises due to covariate-dependent labeling. The method ensures consistency if either the outcome or propensity score model is correctly specified, and achieves a non-standard asymptotic rate dependent on the smaller labeled sample size, enabling valid inference even under model misspecification and imbalanced data.

ABSTRACT

Semi-supervised (SS) inference has received much attention in recent years. Apart from a moderate-sized labeled data, L, the SS setting is characterized by an additional, much larger sized, unlabeled data, U. The setting of |U| >> |L|, makes SS inference unique and different from the standard missing data problems, owing to natural violation of the so-called "positivity" or "overlap" assumption. However, most of the SS literature implicitly assumes L and U to be equally distributed, i.e., no selection bias in the labeling. Inferential challenges in missing at random (MAR) type labeling allowing for selection bias, are inevitably exacerbated by the decaying nature of the propensity score (PS). We address this gap for a prototype problem, the estimation of the response's mean. We propose a double robust SS (DRSS) mean estimator and give a complete characterization of its asymptotic properties. The proposed estimator is consistent as long as either the outcome or the PS model is correctly specified. When both models are correctly specified, we provide inference results with a non-standard consistency rate that depends on the smaller size |L|. The results are also extended to causal inference with imbalanced treatment groups. Further, we provide several novel choices of models and estimators of the decaying PS, including a novel offset logistic model and a stratified labeling model. We present their properties under both high and low dimensional settings. These may be of independent interest. Lastly, we present extensive simulations and also a real data application.

Motivation & Objective

  • To address selection bias in semi-supervised inference when labeled data are small and labeled/unlabeled distributions differ due to covariate-dependent labeling.
  • To develop a double robust estimator that remains consistent if either the outcome regression or propensity score model is correctly specified.
  • To characterize the asymptotic properties of the estimator under decaying overlap, where the propensity score tends to zero as sample size increases.
  • To extend the framework to causal inference with imbalanced treatment groups.
  • To propose novel models for decaying propensity scores, including an offset logistic model and stratified labeling model, under high and low-dimensional settings.

Proposed method

  • Proposes a double robust semi-supervised (DRSS) estimator that combines inverse probability weighting and outcome regression, leveraging both labeled and unlabeled data.
  • Uses a novel offset logistic model to estimate the decaying propensity score, accommodating the diminishing labeling probability in large unlabeled samples.
  • Applies a stratified labeling model to handle heterogeneous labeling mechanisms across subpopulations.
  • Employs a jackknife-type variance estimator to account for the complex dependence induced by the DRSS structure and decaying overlap.
  • Derives asymptotic normality and consistency under weak regularity conditions, with convergence rates depending on the labeled sample size when both models are correct.
  • Uses a cross-fitting procedure with K-fold partitioning to ensure double robustness and avoid overfitting in high-dimensional settings.

Experimental results

Research questions

  • RQ1Can a double robust estimator maintain consistency and validity for mean estimation in semi-supervised settings with MAR labeling and decaying overlap?
  • RQ2What is the asymptotic distribution and convergence rate of the DRSS estimator when the propensity score decays to zero as the sample size grows?
  • RQ3How do model misspecifications in either the outcome regression or propensity score model affect the finite-sample performance of the DRSS estimator?
  • RQ4Can the proposed offset logistic model effectively estimate decaying propensity scores in high-dimensional settings?
  • RQ5How does the DRSS framework extend to causal inference with imbalanced treatment groups under similar selection bias conditions?

Key findings

  • The DRSS estimator is consistent if either the outcome regression or the propensity score model is correctly specified, ensuring robustness to model misspecification.
  • When both models are correctly specified, the estimator achieves a non-standard asymptotic variance that depends on the size of the labeled sample, not the full sample size.
  • The asymptotic distribution is normal with a variance that scales as $ O((n \bar{\pi}_N)^{-1}) $, where $ n $ is the labeled sample size and $ \bar{\pi}_N $ is the average propensity score.
  • The proposed offset logistic model for the propensity score provides consistent estimation even under decaying overlap, outperforming standard logistic regression in simulation studies.
  • The stratified labeling model improves estimation efficiency in settings with heterogeneous labeling mechanisms across subpopulations.
  • Empirical results show that the DRSS estimator maintains valid coverage in finite samples and outperforms standard supervised and semi-supervised estimators under selection bias.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.