Skip to main content
QUICK REVIEW

[论文解读] Doubly Robust Covariate Shift Regression with Semi-nonparametric Nuisance Models

Molei Liu, Yi Zhang|arXiv (Cornell University)|Oct 6, 2020
Statistical Methods and Inference参考文献 58被引用 4
一句话总结

本文提出了一种双重稳健的协变量偏移回归方法,通过将重要性加权与响应变量的半非参数插补模型相结合,降低了对模型误设的敏感性以及维度灾难的影响。当至少一个干扰模型正确时,该方法可确保根n一致性,并实现速率双重稳健的非参数估计,在模拟实验和双相情感障碍表型识别的实际迁移学习任务中,优于参数方法和完全非参数方法。

ABSTRACT

In contemporary statistical learning, covariate shift correction plays an important role when distribution of the testing data is shifted from the training data. Importance weighting is used to adjust for this but is not robust to model misspecifcation or excessive estimation error. In this paper, we propose a doubly robust covariate shift regression approach that introduces an imputation model for the targeted response, and uses it to augment the importance weighting equation. With a novel semi-nonparametric construction for the two nuisance models, our method is less prone to the curse of dimensionality compared to the nonparametric approaches, and is less prone to model mis-specification than the parametric approach. To remove the overfitting bias of the nonparametric components under potential model mis-specification, we construct calibrated moment estimating equations for the semi-nonparametric models. We show that our estimator is root-n consistent when at least one nuisance model is correctly specified, estimation for the parametric part of the nuisance models achieves parametric rate, and the nonparametric components are rate doubly robust. Simulation studies demonstrate that our method is more robust and efficient than existing parametric and fully nonparametric (machine learning) estimators under various configurations. We also examine the utility of our method through a real example about transfer learning of phenotyping algorithm for bipolar disorder. Finally, we propose ways to improve the (intrinsic) efficiency of our estimator and to incorporate high dimensional or machine learning models with our proposed framework.

研究动机与目标

  • 解决测试数据分布与训练数据分布不同的协变量偏移问题。
  • 克服重要性加权方法对模型误设和估计误差敏感的局限性。
  • 在保持对模型误设的稳健性的同时,减少完全非参数方法常见的维度灾难问题。
  • 开发一种即使两个干扰模型(重要性或插补)中有一个被误设仍能保持一致性的方法。
  • 提高估计效率,并实现与高维或机器学习模型的集成。

提出的方法

  • 通过在重要性加权方程中引入响应变量的半非参数插补模型,提出一种双重稳健估计器。
  • 为重要性和插补模型设计了一种新颖的半非参数构造方法,结合了参数与非参数成分。
  • 构建校准的矩估计方程,以在潜在模型误设下纠正非参数成分的过拟合偏差。
  • 在双重稳健性条件下确保根n一致性:至少一个干扰模型被正确指定。
  • 实现参数部分的参数化速率估计,以及非参数成分的速率双重稳健收敛。
  • 提供一个框架,将高维或机器学习模型整合到干扰估计过程中。

实验结果

研究问题

  • RQ1半非参数方法是否能在保持对模型误设的稳健性的同时,减少协变量偏移校正中的维度灾难?
  • RQ2当仅两个干扰模型中的一个被正确指定时,所提出的方法是否能实现根n一致性?
  • RQ3校准的矩估计如何在模型误设下改善非参数成分的偏差校正?
  • RQ4该方法在各种数据配置下是否能优于现有的参数和完全非参数估计器,在稳健性和效率方面表现更优?
  • RQ5如何有效将高维或机器学习模型整合到干扰模型估计框架中?

主要发现

  • 当两个干扰模型(重要性或插补)中至少一个被正确指定时,所提出的估计器可实现根n一致性。
  • 干扰模型中参数部分的估计达到了参数化速率的收敛速度。
  • 干扰模型中非参数部分的收敛速率具有速率双重稳健性,即即使一个模型被误设,其收敛速率也不会受损。
  • 模拟研究显示,该方法在各种分布偏移和模型误设条件下,比参数和完全非参数估计器更具稳健性和高效性。
  • 在双相情感障碍表型识别的真实迁移学习应用中,该方法表现出优于基线方法的性能。
  • 本文提出了对估计器内在效率的高效改进,并为将机器学习模型整合到该框架中提供了可行路径。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。