[论文解读] Semi-supervised Inference: General Theory and Estimation of Means
本文提出了一种半监督推理框架,用于在利用标记响应和未标记协变量的情况下估计总体均值。该方法提出了一种基于最小二乘的估计器,利用协变量分布提高估计效率,即使在中等样本量下也能实现渐近改进,并在最小建模假设下建立了宽度更小的有效置信区间。
We propose a general semi-supervised inference framework focused on the estimation of the population mean. As usual in semi-supervised settings, there exists an unlabeled sample of covariate vectors and a labeled sample consisting of covariate vectors along with real-valued responses ("labels"). Otherwise, the formulation is "assumption-lean" in that no major conditions are imposed on the statistical or functional form of the data. We consider both the ideal semi-supervised setting where infinitely many unlabeled samples are available, as well as the ordinary semi-supervised setting in which only a finite number of unlabeled samples is available. Estimators are proposed along with corresponding confidence intervals for the population mean. Theoretical analysis on both the asymptotic distribution and $\ell_2$-risk for the proposed procedures are given. Surprisingly, the proposed estimators, based on a simple form of the least squares method, outperform the ordinary sample mean. The simple, transparent form of the estimator lends confidence to the perception that its asymptotic improvement over the ordinary sample mean also nearly holds even for moderate size samples. The method is further extended to a nonparametric setting, in which the oracle rate can be achieved asymptotically. The proposed estimators are further illustrated by simulation studies and a real data example involving estimation of the homeless population.
研究动机与目标
- 开发一种无需强参数假设的一般性半监督推断框架,用于估计总体均值。
- 通过在均值估计中整合未标记协变量数据,提高估计效率。
- 在理想(无限未标记数据)和普通(有限未标记数据)设置下,推导所提出估计器的渐近分布和置信区间。
- 证明所提出的估计器在 $\ell_2$-风险和渐近方差方面优于普通样本均值。
- 将该方法扩展至非参数设置,使估计器在渐近下可达到“oracle”率。
提出的方法
- 提出理想半监督估计器 $\hat{\theta} = \bar{\mathbf{Y}} - \hat{\beta}_{(2)}^\top(\bar{\mathbf{X}} - \mu)$,其中 $\mu = \mathbb{E}X$ 假设已知。
- 在有限未标记样本情况下,使用标记和未标记 $X$ 的合并样本均值 $\hat{\mu}$,代入估计器 $\hat{\theta} = \bar{\mathbf{Y}} - \hat{\beta}_{(2)}^\top(\bar{\mathbf{X}} - \hat{\mu})$。
- 采用最小二乘法估计 $\hat{\beta}_{(2)}$,以建模 $Y$ 与 $X$ 之间的关系,而无需假设 $\mathbb{E}(Y|X)$ 的线性关系。
- 在 $p = o(n^{1/2})$ 条件下,推导估计器的渐近分布和 $\ell_2$-风险界。
- 利用矩阵逆展开和集中不等式分析估计器的偏差与方差分量。
- 构建渐近有效的置信区间,其宽度小于基于普通样本均值的区间。
实验结果
研究问题
- RQ1是否可以利用未标记协变量数据在半监督设置下提高均值估计的效率?
- RQ2所提出的估计器在渐近方差和 $\ell_2$-风险方面与普通样本均值相比如何?
- RQ3在理想和有限未标记样本两种情形下,该估计器的理论性质是什么?
- RQ4该方法在非参数设置下是否可达到非参数“oracle”率?
- RQ5能否构建出比基于样本均值更短的置信区间?
主要发现
- 所提出的半监督估计器即使在不假设 $\mathbb{E}(Y|X)$ 线性关系的情况下,其渐近方差也严格小于普通样本均值。
- 估计器的 $\ell_2$-风险严格小于样本均值,表明估计效率得到提升。
- 在已知 $P_X$ 的理想情形下,估计器达到极限分布,从而可构造出比标准方法更短的置信区间。
- 在有限未标记样本情形下,估计器仍保持渐近有效性,并维持更高的估计效率,且在 $p = o(n^{1/2})$ 条件下具有理论保证。
- 在非参数回归设置下,该方法可达到非参数“oracle”率,表明其能最优适应未知光滑度。
- 模拟研究和对无家可归者人口估计的真实数据示例均证实,该方法在实际应用中优于样本均值。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。