Skip to main content
QUICK REVIEW

[论文解读] Double Robust Semi-Supervised Inference for the Mean: Selection Bias under MAR Labeling with Decaying Overlap

Yuqian Zhang, Abhishek Chakrabortty|arXiv (Cornell University)|Apr 14, 2021
Advanced Causal Inference Techniques参考文献 39被引用 4
一句话总结

本文提出了一种在缺失 at random (MAR) 标注且重叠度衰减条件下的双 robust 半监督 (DRSS) 估计量,用于均值估计,其中标注数据稀少,且由于协变量依赖的标注机制导致选择偏差。该方法在结果模型或倾向得分模型任一正确指定时均能保持一致性,并实现一种非标准的渐近收敛速率,该速率依赖于较小的标注样本量,从而在模型误设和数据不平衡的情况下仍能实现有效的推断。

ABSTRACT

Semi-supervised (SS) inference has received much attention in recent years. Apart from a moderate-sized labeled data, L, the SS setting is characterized by an additional, much larger sized, unlabeled data, U. The setting of |U| >> |L|, makes SS inference unique and different from the standard missing data problems, owing to natural violation of the so-called "positivity" or "overlap" assumption. However, most of the SS literature implicitly assumes L and U to be equally distributed, i.e., no selection bias in the labeling. Inferential challenges in missing at random (MAR) type labeling allowing for selection bias, are inevitably exacerbated by the decaying nature of the propensity score (PS). We address this gap for a prototype problem, the estimation of the response's mean. We propose a double robust SS (DRSS) mean estimator and give a complete characterization of its asymptotic properties. The proposed estimator is consistent as long as either the outcome or the PS model is correctly specified. When both models are correctly specified, we provide inference results with a non-standard consistency rate that depends on the smaller size |L|. The results are also extended to causal inference with imbalanced treatment groups. Further, we provide several novel choices of models and estimators of the decaying PS, including a novel offset logistic model and a stratified labeling model. We present their properties under both high and low dimensional settings. These may be of independent interest. Lastly, we present extensive simulations and also a real data application.

研究动机与目标

  • 解决在标注数据稀少且标注/未标注分布因协变量依赖的标注机制而不同的半监督推断中的选择偏差问题。
  • 开发一种双 robust 估计量,当结果回归模型或倾向得分模型任一正确指定时,仍能保持一致性。
  • 刻画在重叠度衰减条件下估计量的渐近性质,其中倾向得分随样本量增加趋于零。
  • 将该框架扩展至处理治疗组不平衡的因果推断问题。
  • 在高维和低维设定下,提出新型的衰减倾向得分模型,包括偏移逻辑斯蒂克模型和分层标注模型。

提出的方法

  • 提出一种双 robust 半监督 (DRSS) 估计量,结合逆概率加权和结果回归,利用标注和未标注数据。
  • 使用一种新颖的偏移逻辑斯蒂克模型来估计衰减的倾向得分,以适应在大规模未标注样本中标签概率递减的情形。
  • 应用分层标注模型以处理不同子群体间标注机制异质性的问题。
  • 采用一种类似刀切法的方差估计量,以考虑 DRSS 结构和衰减重叠所引起的复杂依赖性。
  • 在弱正则性条件下推导出渐近正态性和一致性,当两个模型均正确时,收敛速率依赖于标注样本量。
  • 在高维设定下使用 K 折划分的交叉拟合程序,以确保双 robust 性并避免过拟合。

实验结果

研究问题

  • RQ1在 MAR 标注且重叠度衰减的半监督设定下,双 robust 估计量能否保持一致性与有效性?
  • RQ2当倾向得分随样本量增长趋于零时,DRSS 估计量的渐近分布与收敛速率为何?
  • RQ3结果回归或倾向得分模型中的模型误设如何影响 DRSS 估计量的有限样本性能?
  • RQ4所提出的偏移逻辑斯蒂克模型在高维设定下能否有效估计衰减的倾向得分?
  • RQ5在类似的选择偏差条件下,DRSS 框架如何扩展至治疗组不平衡的因果推断?

主要发现

  • 当结果回归模型或倾向得分模型任一正确指定时,DRSS 估计量具有一致性,确保对模型误设的鲁棒性。
  • 当两个模型均正确时,估计量达到一种非标准的渐近方差,其依赖于标注样本量,而非完整样本量。
  • 渐近分布为正态分布,其方差按 $ O((n \bar{\pi}_N)^{-1}) $ 的尺度缩放,其中 $ n $ 为标注样本量,$ \bar{\pi}_N $ 为平均倾向得分。
  • 所提出的倾向得分偏移逻辑斯蒂克模型在重叠度衰减条件下仍能提供一致估计,在模拟研究中优于标准逻辑斯蒂克回归。
  • 分层标注模型在子群体间标注机制异质性较高的设定下,可提升估计效率。
  • 实证结果表明,DRSS 估计量在有限样本中保持了有效的置信区间覆盖,并在选择偏差条件下优于标准监督和半监督估计量。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。