Skip to main content
QUICK REVIEW

[论文解读] On the multiply robust estimation of the mean of the g-functional

Andrea Rotnitzky, James M. Robins|arXiv (Cornell University)|May 24, 2017
Statistical Methods and Inference参考文献 7被引用 19
一句话总结

本文提出了一种基于机器学习的纵向g-函数多重稳健(MR)估计量,通过样本分割避免了Donsker条件,展示了在大多数数据生成机制下,MR估计量与双重稳健(DR)估计量具有渐近等价的收敛速度。尽管在特定数据生成规律下MR估计量可能比DR估计量收敛更快,但通常两者收敛速率相同,MR估计量在多数情形下并无一致优势。

ABSTRACT

We study multiply robust (MR) estimators of the longitudinal g-computation formula of Robins (1986). In the first part of this paper we review and extend the recently proposed parametric multiply robust estimators of Tchetgen-Tchetgen (2009) and Molina, Rotnitzky, Sued and Robins (2017). In the second part of the paper we derive multiply and doubly robust estimators that use non-parametric machine-learning (ML) estimators of nuisance functions in lieu of parametric models. We use sample splitting to avoid the need for Donsker conditions, thereby allowing an analyst to select the ML algorithms of their choosing. We contrast the asymptotic behavior of our non-parametric doubly robust and multiply robust estimators. In particular, we derive formulas for their asymptotic bias. Examining these formulas we conclude that although, under certain data generating laws, the rate at which the bias of the MR estimator converges to zero can exceed that of the DR estimator, nonetheless, under most laws, the bias of the DR and MR estimators converge to zero at the same rate.

研究动机与目标

  • 将参数化的多重稳健估计量扩展至非参数设置,利用机器学习估计干扰函数。
  • 通过样本分割避免Donsker条件,基于最小建模假设,发展g-函数的渐近有效估计量。
  • 比较使用机器学习算法时,多重稳健与双重稳健估计量的渐近偏差率。
  • 探究在特定数据生成机制下,多重稳健估计量是否可实现比双重稳健估计量更快的收敛速率。

提出的方法

  • 使用样本分割构造交叉拟合估计量,确保无需对机器学习算法施加Donsker型条件即可保持有效性。
  • 推导在一般机器学习算法下,双重稳健(DR)与多重稳健(MR)估计量的渐近偏差公式。
  • 利用逆概率加权和g-函数的迭代回归表示,构建MR估计量。
  • 提出两类MR估计量:K+1重稳健与2^K重稳健,基于不同的建模假设。
  • 通过将渐近偏差分解为机器学习估计残差相关项,分析DR与MR估计量的偏差漂移。
  • 通过评估渐近展开中特定偏差分量的主导性,比较DR与MR估计量的收敛速率。

实验结果

研究问题

  • RQ1在何种数据生成规律下,多重稳健估计量可比双重稳健估计量更快收敛至真实参数?
  • RQ2在采用样本分割时,使用机器学习估计干扰函数如何影响g-函数估计量的渐近偏差?
  • RQ3不同偏差分量(如ρ_k^MR与χ_j,k^DR)对MR与DR估计量整体偏差漂移的相对贡献为何?
  • RQ42^K重稳健估计量中模型组合数量的增加是否显著提升其比DR估计量更快收敛的概率?
  • RQ5在一般机器学习算法下,MR估计量的渐近偏差是否能保证以与DR估计量相同的速度收敛至零?

主要发现

  • 在大多数数据生成规律下,多重稳健估计量的渐近偏差收敛至零的速度与双重稳健估计量相同。
  • 在某些特定数据生成规律下,当DR估计量的偏差漂移中某些交叉项(χ_j,k^DR)占主导时,MR估计量可实现更快收敛。
  • ρ_k^MR与ρ_k^DR偏差分量的收敛速率通常由收敛最慢的残差项决定,该残差项通常涉及最终时间点。
  • DR估计量偏差漂移中潜在主导项的数量随K的平方增长,使得χ_j,k^DR在K较大时可能主导ρ^DR。
  • 对于非线性机器学习算法,无法保证R_MR,k残差比R_DR,k收敛得更快,从而限制了MR估计量的潜在优势。
  • 样本分割确保了渐近理论的有效性,无需满足Donsker条件,从而可灵活使用任意机器学习算法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。