Skip to main content
QUICK REVIEW

[论文解读] Finite Sample Inference for Targeted Learning

Mark van der Laan|arXiv (Cornell University)|Aug 30, 2017
Advanced Causal Inference Techniques参考文献 20被引用 3
一句话总结

本文提出了四种基于有限样本推断的方法,用于目标最大似然估计(TMLE)与高度自适应Lasso(HAL)结合,利用非参数自展法估计HAL-TMLE的真实抽样分布,同时考虑高阶余项的影响。关键贡献在于提出了一种理论上有依据、渐近一致的自展法,相较于基于正态近似的推断,显著提升了在高维或复杂模型下的有限样本覆盖性能。

ABSTRACT

The Highly-Adaptive-Lasso(HAL)-TMLE is an efficient estimator of a pathwise differentiable parameter in a statistical model that at minimal (and possibly only) assumes that the sectional variation norm of the true nuisance parameters are finite. It relies on an initial estimator (HAL-MLE) of the nuisance parameters by minimizing the empirical risk over the parameter space under the constraint that sectional variation norm is bounded by a constant, where this constant can be selected with cross-validation. In the formulation of the HALMLE this sectional variation norm corresponds with the sum of absolute value of coefficients for an indicator basis. Due to its reliance on machine learning, statistical inference for the TMLE has been based on its normal limit distribution, thereby potentially ignoring a large second order remainder in finite samples. In this article, we present four methods for construction of a finite sample 0.95-confidence interval that use the nonparametric bootstrap to estimate the finite sample distribution of the HAL-TMLE or a conservative distribution dominating the true finite sample distribution. We prove that it consistently estimates the optimal normal limit distribution, while its approximation error is driven by the performance of the bootstrap for a well behaved empirical process. We demonstrate our general inferential methods for 1) nonparametric estimation of the average treatment effect based on observing on each unit a covariate vector, binary treatment, and outcome, and for 2) nonparametric estimation of the integral of the square of the multivariate density of the data distribution.

研究动机与目标

  • 为解决基于正态近似的TMLE推断的局限性,该方法在有限样本中可能因高阶余项过大而产生不准确结果。
  • 为HAL-TMLE开发对模型复杂性和高维数据具有鲁棒性的有限样本置信区间。
  • 在弱正则性条件下,建立非参数自展法估计HAL-TMLE真实抽样分布的理论有效性。
  • 证明自展法能够保持HAL-MLE的渐近行为,并相较于Wald型区间提升推断精度。
  • 探讨截面变差范数作为稀疏性度量的作用,及其对推断性能的影响。

提出的方法

  • 本文使用非参数自展法估计HAL-TMLe的有限样本分布,将交叉验证得到的截面变差范数界视为固定值。
  • 提出四种基于自展法的推断方法:两种用于估计真实抽样分布,两种为保守变体,其分布严格主导真实分布。
  • 该方法依赖HAL-MLE来估计具有有界截面变差范数的异质性参数,并通过交叉验证选择。
  • 利用目标参数的自然梯度(影响曲线)定义HAL-TMLE的渐近线性展开。
  • 将自展法应用于HAL-TMLE的精确二阶展开,以直接估计其抽样分布,尽管这会损失替代估计器的性质。
  • 在弱正则性条件下,证明了自展法的理论一致性,近似误差由自展法在表现良好的经验过程上的表现决定。

实验结果

研究问题

  • RQ1当高阶余项较大时,非参数自展法能否一致地估计HAL-TMLE的真实有限样本分布?
  • RQ2通过交叉验证选择的截面变差范数界如何影响HAL-TMLE推断的有限样本性能?
  • RQ3基于自展法的推断是否优于依赖渐近正态性的标准Wald型置信区间?
  • RQ4截面变差范数作为稀疏性度量,在实现高维非参数模型中稳健自适应推断方面发挥何种作用?
  • RQ5自展法能否扩展至纵向因果推断中使用的序列或递归HAL-TMLE?

主要发现

  • 非参数自展法能一致地估计HAL-TMLE最优正态极限分布,近似误差由自展法在经验过程上的表现决定。
  • 自展法保持了HAL-MLE在异质性参数上的渐近行为,为其在推断中的应用提供了理论支持。
  • 所提出的保守自展法严格主导真实有限样本分布,即使真实截面变差范数被低估,也能确保有效覆盖。
  • 基于自展法的有限样本置信区间比基于正态近似的区间更准确,尤其在高维设置下、异质性参数空间较大的情况下。
  • 截面变差范数是一种强大且与基无关的稀疏性度量,可在弱正则性条件下实现自适应、稳健的估计与推断。
  • 结果表明,通过无限个指示函数的线性组合表示函数的指标基表示,是定义复杂性的唯一合适方式,可实现高效的MLE并支持有效的自展推断。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。