Skip to main content
QUICK REVIEW

[论文解读] Higher Order Targeted Maximum Likelihood Estimation

Mark van der Laan, Zeyi Wang|arXiv (Cornell University)|Jan 15, 2021
Statistical Methods and Inference参考文献 26被引用 4
一句话总结

本文提出了一种高阶目标似然估计量(k阶TMLE),通过用k+1阶余项替代标准TMLE中的首阶余项,结合高度自适应Lasso(HAL)正则化以控制偏差,从而改进小样本推断。该方法在较弱的正则性条件下确保渐近线性并实现有效推断,模拟结果证实其在覆盖区间和偏差减少方面优于首阶TMLE。

ABSTRACT

Asymptotic efficiency of targeted maximum likelihood estimators (TMLE) of target features of the data distribution relies on a a second order remainder being asymptotically negligible. In previous work we proposed a nonparametric MLE termed Highly Adaptive Lasso (HAL) which parametrizes the relevant functional of the data distribution in terms of a multivariate real valued cadlag function that is assumed to have finite variation norm. We showed that the HAL-MLE converges in Kullback-Leibler dissimilarity at a rate n-1/3 up till logn factors. Therefore, by using HAL as initial density estimator in the TMLE, the resulting HAL-TMLE is an asymptotically efficient estimator only assuming that the relevant nuisance functions of the data density are cadlag and have finite variation norm. However, in finite samples, the second order remainder can dominate the sampling distribution so that inference based on asymptotic normality would be anti-conservative. In this article we propose a new higher order TMLE, generalizing the regular first order TMLE. We prove that it satisfies an exact linear expansion, in terms of efficient influence functions of sequentially defined higher order fluctuations of the target parameter, with a remainder that is a k+1th order remainder. As a consequence, this k-th order TMLE allows statistical inference only relying on the k+1th order remainder being negligible. We also provide a rationale for the higher order TMLE that it will be superior to the first order TMLE by (iteratively) locally minimizing the exact finite sample remainder of the first order TMLE. The second order TMLE is demonstrated for nonparametric estimation of the integrated squared density and for the treatment specific mean outcome. We also provide an initial simulation study for the second order TMLE of the treatment specific mean confirming the theoretical analysis.

研究动机与目标

  • 解决由非可忽略的二阶余项引起的靶向最大似然估计(TMLE)在小样本中的偏差问题。
  • 构建一种k阶TMLE框架,确保其具有精确的线性展开形式,余项为k+1阶,从而在无需欠平滑的情况下实现有效推断。
  • 证明通过控制HAL-MLE的L1-范数,HAL正则化的高阶TMLE即使在无欠平滑条件下,其正则化偏差也可忽略不计。
  • 提供小样本合理性论证,表明高阶TMLE通过在初始估计量周围迭代地最小化首阶TMLE的精确余项,实现偏差控制。
  • 在首阶影响函数在真值处消失的情况下,利用二阶影响函数实现推断。

提出的方法

  • 提出一种k阶TMLE,通过沿最不利路径使用HAL正则化的MLE,依次靶向目标参数的高阶波动。
  • 采用通用的最不利路径构造方法,确保每一步的得分等于波动后参数的首阶影响函数。
  • 推导出k阶TMLE在高阶波动的效率影响函数中的精确展开形式,余项阶数为k+1。
  • 通过控制HAL-MLE的L1-范数来控制HAL正则化偏差,确保其在无欠平滑条件下渐近可忽略。
  • 利用经验过程理论与熵界,证明在最小平滑性假设下,k+1阶余项可忽略不计。
  • 提供一种基于伴随算子与对称矩阵求逆的构造性算法,可系统化计算高阶影响函数,支持使用标准软件实现。

实验结果

研究问题

  • RQ1能否构造一种k阶TMLE,使其余项为k+1阶,从而在不依赖欠平滑的情况下实现有效推断?
  • RQ2对扰动估计量进行HAL正则化是否能确保其正则化偏差在小样本中可忽略?
  • RQ3与首阶TMLE相比,高阶TMLE是否能在小样本中提升覆盖区间并减少偏差?
  • RQ4当首阶影响函数在真值处为零时,是否可利用二阶影响函数实现推断?
  • RQ5能否使用标准回归或矩阵求逆工具系统化计算高阶影响函数?

主要发现

  • k阶TMLE满足精确的线性展开,余项阶数为k+1,因此推断仅依赖于该余项的可忽略性。
  • 由于对HAL-MLE的L1-范数进行了控制,HAL正则化偏差在无欠平滑条件下被保证为足够小的阶数。
  • 对积分平方密度和处理特定均值结果的非参数估计的模拟结果表明,二阶TMLE相比首阶TMLE显著降低了偏差并改善了置信区间覆盖。
  • 经验高阶TMLE的性能与欠平滑版本相当,表明其具有稳健性,是推荐的实用选择。
  • 当首阶影响函数在真值处为零时,该方法仍可通过使用二阶影响函数获得非退化的极限分布,从而实现推断。
  • 提供了一种基于伴随算子与矩阵求逆的构造性算法,可系统化计算高阶影响函数,支持使用标准软件实现。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。