[论文解读] Semi-Supervised Quantile Estimation: Robust and Efficient Inference in High Dimensional Settings
本文提出了一种半监督分位数估计方法,利用大规模未标记数据集在高维设置下提升响应分位数估计的准确性与推理效率。通过结合模型无关的插补方法、去偏步骤及一步更新,该方法即使在插补模型误设的情况下,仍能实现根n一致性与渐近正态性,而在模型正确设定时可达到半参数效率。
We consider quantile estimation in a semi-supervised setting, characterized by two available data sets: (i) a small or moderate sized labeled data set containing observations for a response and a set of possibly high dimensional covariates, and (ii) a much larger unlabeled data set where only the covariates are observed. We propose a family of semi-supervised estimators for the response quantile(s) based on the two data sets, to improve the estimation accuracy compared to the supervised estimator, i.e., the sample quantile from the labeled data. These estimators use a flexible imputation strategy applied to the estimating equation along with a debiasing step that allows for full robustness against misspecification of the imputation model. Further, a one-step update strategy is adopted to enable easy implementation of our method and handle the complexity from the non-linear nature of the quantile estimating equation. Under mild assumptions, our estimators are fully robust to the choice of the nuisance imputation model, in the sense of always maintaining root-n consistency and asymptotic normality, while having improved efficiency relative to the supervised estimator. They also attain semi-parametric optimality if the relation between the response and the covariates is correctly specified via the imputation model. As an illustration of estimating the nuisance imputation function, we consider kernel smoothing type estimators on lower dimensional and possibly estimated transformations of the high dimensional covariates, and we establish novel results on their uniform convergence rates in high dimensions, involving responses indexed by a function class and usage of dimension reduction techniques. These results may be of independent interest. Numerical results on both simulated and real data confirm our semi-supervised approach's improved performance, in terms of both estimation and inference.
研究动机与目标
- 解决现代生物医学与观察性研究中因标记数据有限而导致高维分位数估计统计功效不足的问题。
- 开发一种半监督推理框架,利用丰富的未标记协变量提升估计效率与鲁棒性。
- 确保即使在插补模型误设的情况下,分位数估计量仍保持根n一致性与渐近正态性。
- 在模型正确设定时实现半参数效率,从而在精度上优于监督估计方法。
- 建立适用于扰动函数估计的高维降维下核平滑的新型一致收敛速率。
提出的方法
- 提出一类基于灵活插补策略应用于分位数估计方程的半监督估计量族。
- 引入去偏步骤以确保对插补模型误设的鲁棒性,同时保持根n一致性与渐近正态性。
- 采用一步更新以简化实现过程,并处理分位数估计方程的非线性特性。
- 在高维协变量的低维投影上使用核平滑估计扰动插补函数。
- 应用降维技术(如线性回归、切片逆回归)在核平滑前降低有效维度。
- 建立函数类索引响应下高维设置中核平滑估计量的新一致收敛速率。
实验结果
研究问题
- RQ1在标记数据有限的高维设置下,能否利用未标记数据提升分位数估计的准确性?
- RQ2在半监督分位数估计中,如何确保对插补模型误设的鲁棒性?
- RQ3在模型误设下,半监督分位数估计量的渐近行为如何?
- RQ4在模型正确设定时,能否在分位数估计中实现半参数效率?
- RQ5当应用于高维降维协变量时,核平滑估计量的统一收敛速率是什么?
主要发现
- 所提估计量在任意插补模型下均保持根n一致性与渐近正态性,确保对模型误设的鲁棒性。
- 与监督样本分位数相比,估计量效率得到提升,在模拟中相对效率增益最高达30%。
- 当插补模型正确设定响应条件均值时,可实现半参数效率。
- 在降维协变量上进行核平滑可实现随高维输入呈有利趋势的统一收敛速率。
- 数值结果表明,与监督方法相比,该半监督方法能获得更精确的分位数估计,并实现更优的95%置信区间覆盖。
- 在真实数据分析中,该方法在具有高维协变量的大规模健康调查数据集中展现出更优的推断性能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。