Skip to main content
QUICK REVIEW

[论文解读] Nonparametric Bootstrap Inference for the Targeted Highly Adaptive LASSO Estimator

Weixin Cai, Mark van der Laan|arXiv (Cornell University)|May 23, 2019
Statistical Methods and Inference参考文献 19被引用 5
一句话总结

本文提出了一种针对目标高度自适应LASSO估计量(HAL-TMLE)的非参数自展法,通过将分段变差范数固定在大于或等于交叉验证选择器的值,确保渐近一致性。该方法引入了一种基于平台的范数选择器,以优化小样本置信区间覆盖,模拟结果显示其在平均处理效应和密度积分估计方面优于Wald型区间。

ABSTRACT

The Highly-Adaptive-LASSO Targeted Minimum Loss Estimator (HAL-TMLE) is an efficient plug-in estimator of a pathwise differentiable parameter in a statistical model that at minimal (and possibly only) assumes that the sectional variation norm of the true nuisance functional parameters (i.e., the relevant part of data distribution) are finite. It relies on an initial estimator (HAL-MLE) of the nuisance functional parameters by minimizing the empirical risk over the parameter space under the constraint that the sectional variation norm of the candidate functions are bounded by a constant, where this constant can be selected with cross-validation. In this article, we establish that the nonparametric bootstrap for the HAL-TMLE, fixing the value of the sectional variation norm at a value larger or equal than the cross-validation selector, provides a consistent method for estimating the normal limit distribution of the HAL-TMLE. In order to optimize the finite sample coverage of the nonparametric bootstrap confidence intervals, we propose a selection method for this sectional variation norm that is based on running the nonparametric bootstrap for all values of the sectional variation norm larger than the one selected by cross-validation, and subsequently determining a value at which the width of the resulting confidence intervals reaches a plateau. We demonstrate our method for 1) nonparametric estimation of the average treatment effect based on observing on each unit a covariate vector, binary treatment, and outcome, and for 2) nonparametric estimation of the integral of the square of the multivariate density of the data distribution. In addition, we also present simulation results for these two examples demonstrating the excellent finite sample coverage of bootstrap-based confidence intervals.

研究动机与目标

  • 为解决HAL-TMLE在Wald型置信区间中因高阶余项过大而导致的有限样本覆盖性能差的问题,该问题可能具有反保守性。
  • 开发一种自展方法,通过将分段变差范数固定在数据自适应的值,一致地估计HAL-TMLE的真实抽样分布。
  • 通过基于平台检测的方法选择分段变差范数,对自展法进行有限样本优化,以提升区间覆盖性能。
  • 在高维设定下,展示该方法在非参数平均处理效应和多元密度积分估计中的有效性。

提出的方法

  • 非参数自展法将HAL-MLE的分段变差范数固定在一个大于或等于交叉验证选择器的值,确保HAL-TMLE的自展分布的一致性。
  • 该方法使用经验分布作为抽样分布,利用HAL-MLE在重采样下即使位于参数空间边界也具有鲁棒性的特性。
  • 提出一种平台选择器:在交叉验证选择值以上的所有分段变差范数值上运行自展,选择置信区间宽度稳定(达到平台)的值作为最终选择。
  • 最终的置信区间通过一个同时考虑自展分布中偏差与方差的方差估计器进行缩放。
  • 该方法应用于两个示例:非参数平均处理效应估计和多元密度平方的积分估计。
  • 该方法继承了HAL-TMLE的渐近效率,并由于HAL-MLE收敛行为良好,确保了重采样下的一致性。

实验结果

研究问题

  • RQ1当分段变差范数被固定在大于或等于交叉验证选择器的值时,非参数自展法能否一致地估计HAL-TMLE的有限样本抽样分布?
  • RQ2如何在自展法中优化分段变差范数的选择,以提升有限样本置信区间的覆盖性能?
  • RQ3所提出的基于平台的分段变差范数选择器是否能带来比直接在自展中使用交叉验证选择器更好的覆盖性能?
  • RQ4在高维、非参数设定下,基于自展的推断与Wald型推断相比,在覆盖性能和区间宽度方面表现如何?

主要发现

  • 当分段变差范数被固定在大于或等于交叉验证选择器的值时,HAL-TMLE的非参数自展法对正态极限分布具有渐近一致性。
  • 基于平台的分段变差范数选择器在平均处理效应和密度积分估计的小样本模拟中,产生了具有优异有限样本覆盖性能的置信区间。
  • 该自展方法通过捕捉更高阶的随机行为,优于Wald型区间,减少了小样本中的反保守性。
  • 即使交叉验证选择器低估了真实的分段变差范数,该方法仍保持稳健,得益于基于平台的优化修正。
  • 自展分布近似误差主要由自展法对行为良好的经验过程的有限样本分布估计的准确性所驱动。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。