[论文解读] Selective machine learning of doubly robust functionals
本文提出了一种选择性机器学习框架,用于估计具有双重稳健性的半参数功能性,通过一种新颖的伪风险准则,最小化由选择处理异质性参数的最优机器学习算法所导致的偏差。该方法通过多折交叉验证实现类似Oracle的性能,并在模拟和真实世界数据中展示了在平均处理效应估计方面改进的稳健性。
While model selection is a well-studied topic in parametric and nonparametric regression or density estimation, selection of possibly high-dimensional nuisance parameters in semiparametric problems is far less developed. In this paper, we propose a selective machine learning framework for making inferences about a finite-dimensional functional defined on a semiparametric model, when the latter admits a doubly robust estimating function and several candidate machine learning algorithms are available for estimating the nuisance parameters. We introduce a new selection criterion aimed at bias reduction in estimating the functional of interest based on a novel definition of pseudo-risk inspired by the double robustness property. Intuitively, the proposed criterion selects a pair of learners with the smallest pseudo-risk, so that the estimated functional is least sensitive to perturbations of a nuisance parameter. We establish an oracle property for a multi-fold cross-validation version of the new selection criterion which states that our empirical criterion performs nearly as well as an oracle with a priori knowledge of the pseudo-risk for each pair of candidate learners. Finally, we apply the approach to model selection of a semiparametric estimator of average treatment effect given an ensemble of candidate machine learners to account for confounding in an observational study which we illustrate in simulations and a data application.
研究动机与目标
- 为解决半参数推断中高维异质性参数缺乏模型选择方法的问题,特别是当感兴趣的泛函具有双重稳健估计函数时。
- 开发一种选择准则,通过最小化对异质性参数估计扰动的敏感性,减少估计泛函的偏差。
- 为所提出的选型准则的交叉验证版本建立理论保证——特别是Oracle性质。
- 将该方法应用于观察性研究中的平均处理效应估计,通过集成机器学习模型控制混杂因素。
- 与不考虑功能敏感性的预设单一学习器的标准方法相比,展示在估计中更高的稳健性与准确性。
提出的方法
- 提出一种基于双重稳健性特性的新型伪风险度量,定义为:在固定一个异质性参数的同时,对另一个异质性参数施加扰动时,泛函估计值的平均平方偏差。
- 采用多折交叉验证方案,估计每对候选学习器(一个用于倾向得分,一个用于结果回归)的伪风险。
- 选择伪风险估计值最小的学习器对,即对任一异质性参数扰动最不敏感的组合。
- 使用样本分割计算泛函估计值与方差,确保渐近正态性与有效推断。
- 将选择准则应用于一大类候选机器学习模型,包括Lasso、带L1正则化的逻辑回归及其他灵活估计器。
- 在选择准则中引入偏差校正项,以考虑任一异质性模型可能存在的模型误设,利用双重稳健结构。
实验结果
研究问题
- RQ1能否开发一种模型选择准则,明确针对双重稳健泛函的偏差减少,而非预测误差?
- RQ2基于伪风险的交叉验证选择规则是否能实现接近Oracle的性能,即若已知每对学习器的真实伪风险?
- RQ3与标准方法(如目标最大似然估计或超级学习器)相比,该方法在平均处理效应估计中的偏差与置信区间覆盖性能如何?
- RQ4该方法能否有效选择对结果模型或倾向得分模型误设具有鲁棒性的机器学习算法?
- RQ5在高维混杂设定下,模型选择对平均处理效应置信区间宽度与准确性的影 响如何?
主要发现
- 基于伪风险的所提选择准则实现了Oracle性质,意味着其性能几乎等同于事先已知真实伪风险的情况。
- 在SUPPORT数据应用中,该方法为倾向得分选择带L1正则化的逻辑回归,为结果模型选择Lasso,估计的平均处理效应为-0.0661。
- 估计的95%置信区间为(-0.0976, -0.0345),略宽于其他改进型估计器,但因更优的模型选择而可能具有更低偏差,因而更准确。
- 该估计效应小于TMLE结合超级学习器的结果(-0.0586)、BR估计器(-0.0610至-0.0612)以及校准似然估计器(-0.0622),表明先前估计值仍可能受小规模模型误设偏差的影响。
- 该方法通过选择最小化对异质性参数扰动敏感性的模型,展现出稳健性,支持其在半参数推断中减少偏差的作用。
- 理论与实证结果共同表明,通过伪风险实现偏差感知的模型选择,相比标准方法,在双重稳健设定下能带来更可靠的推断。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。