[论文解读] A General Framework for Treatment Effect Estimation in Semi-Supervised and High Dimensional Settings
本文提出了一种通用的半监督框架,用于在高维设置下利用标注数据(响应变量和协变量)和未标注数据(仅协变量)来估计平均处理效应和分位数处理效应。通过利用大规模未标注样本,该方法在正确设定倾向得分的前提下实现了根n一致性与渐近正态性,相较于监督方法具有更高的稳健性与效率,即使在高维下对干扰函数进行非参数估计时亦是如此。
In this article, we aim to provide a general and complete understanding of semi-supervised (SS) causal inference for treatment effects. Specifically, we consider two such estimands: (a) the average treatment effect and (b) the quantile treatment effect, as prototype cases, in an SS setting, characterized by two available data sets: (i) a labeled data set of size $n$, providing observations for a response and a set of high dimensional covariates, as well as a binary treatment indicator; and (ii) an unlabeled data set of size $N$, much larger than $n$, but without the response observed. Using these two data sets, we develop a family of SS estimators which are ensured to be: (1) more robust and (2) more efficient than their supervised counterparts based on the labeled data set only. Beyond the 'standard' double robustness results (in terms of consistency) that can be achieved by supervised methods as well, we further establish root-n consistency and asymptotic normality of our SS estimators whenever the propensity score in the model is correctly specified, without requiring specific forms of the nuisance functions involved. Such an improvement of robustness arises from the use of the massive unlabeled data, so it is generally not attainable in a purely supervised setting. In addition, our estimators are shown to be semi-parametrically efficient as long as all the nuisance functions are correctly specified. Moreover, as an illustration of the nuisance estimators, we consider inverse-probability-weighting type kernel smoothing estimators involving unknown covariate transformation mechanisms, and establish in high dimensional scenarios novel results on their uniform convergence rates, which should be of independent interest. Numerical results on both simulated and real data validate the advantage of our methods over their supervised counterparts with respect to both robustness and efficiency.
研究动机与目标
- 开发一种通用框架,用于在半监督、高维设置下估计平均处理效应与分位数处理效应。
- 通过结合大规模未标注数据与小规模标注数据,提升估计效率与稳健性。
- 在正确设定倾向得分的前提下,建立估计量的根n一致性与渐近正态性,且无需对干扰函数的形式作特定假设。
- 推导高维、非参数干扰函数估计器(如具有未知协变量变换的核平滑)的新型一致收敛速率。
- 通过模拟与真实数据分析验证该方法相较于监督方法的优越性。
提出的方法
- 该框架结合标注数据(含结果)与未标注数据(无结果)以构建平均处理效应与分位数处理效应的双重稳健估计量。
- 采用基于逆概率加权的核平滑方法进行干扰函数估计,适用于高维设置下协变量变换未知的情形。
- 当倾向得分正确设定时,该方法可确保估计量的根n一致性与渐近正态性,即使结果回归或其他干扰函数估计不一致亦成立。
- 当所有干扰函数均正确设定时,该方法可实现半参数效率,充分利用大规模未标注样本的信息。
- 该方法采用两阶段估计策略:首先从标注数据中估计干扰函数,然后将结果与未标注数据结合以改进推断。
- 在最小正则性条件下推导理论保证,包括基于核估计器的高维收敛速率。
实验结果
研究问题
- RQ1与纯监督方法相比,半监督框架是否能在高维设置下提升处理效应估计的稳健性与效率?
- RQ2当倾向得分正确设定时,半监督估计量在何种条件下可实现根n一致性与渐近正态性?
- RQ3当干扰函数(如结果回归或倾向得分)在高维下以非参数方式估计时,所提估计量表现如何?
- RQ4在协变量变换未知的高维半监督设置下,逆概率加权核平滑估计器的统一收敛速率为何?
- RQ5在处理效应估计中,引入大规模未标注数据在多大程度上可降低估计偏差与标准误?
主要发现
- 所提出的半监督估计量在正确设定倾向得分的前提下,即使其他干扰函数估计不一致,仍可实现根n一致性与渐近正态性。
- 模拟与真实数据结果表明,该方法相较于监督方法显著提升了效率与稳健性,表现为置信区间更窄与均方误差更低。
- 在HIV药物耐药性数据分析中,半监督置信区间始终短于监督方法,其中最短区间以蓝色突出显示,表明精度提升。
- 对于高维干扰函数估计,本文首次建立了具有未知协变量变换的逆概率加权核平滑估计器的新型一致收敛速率。
- 理论结果表明,当所有干扰函数均正确设定时,估计量可实现半参数效率,充分挖掘未标注数据的全部信息。
- 数值结果表明,该方法在各种模拟设置下均保持良好的覆盖率(约95%),即使结果模型存在模型误设亦成立。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。