[论文解读] Inference for treatment-specific survival curves using machine learning
本文提出了一种基于机器学习估计条件生存函数的交叉拟合双重稳健估计量,用于处理特定治疗的生存曲线,确保在弱正则性条件下的一致性和渐近正态性。该方法可在连续或离散时间下实现有效的推断,并引入了一种新颖的集成学习方法,用于组合多个生存估计量,从而在存在混杂因素的观察性研究中提高稳健性和效率。
In the absence of data from a randomized trial, researchers often aim to use observational data to draw causal inference about the effect of a treatment on a time-to-event outcome. In this context, interest often focuses on the treatment-specific survival curves; that is, the survival curves were the entire population under study to be assigned to receive the treatment or not. Under certain causal conditions, including that all confounders of the treatment-outcome relationship are observed, the treatment-specific survival can be identified with a covariate-adjusted survival function. Several estimators of this function have been proposed, including estimators based on outcome regression, inverse probability weighting, and doubly robust estimators. In this article, we propose a new cross-fitted doubly-robust estimator that incorporates data-adaptive (e.g. machine learning) estimators of the conditional survival functions. We establish conditions on the nuisance estimators under which our estimator is consistent and asymptotically linear, both pointwise and uniformly in time. We also propose a novel ensemble learner for combining multiple candidate estimators of the conditional survival estimators. Notably, our methods and results accommodate events occurring in discrete or continuous time (or both). We investigate the practical performance of our methods using numerical studies and an application to the effect of a surgical treatment to prevent metastases of parotid carcinoma on mortality.
研究动机与目标
- 开发一种在存在混杂因素的观察性研究中,用于处理特定治疗生存曲线的双重稳健估计量。
- 引入数据自适应的机器学习方法来估计条件生存函数,以提高对模型误设的稳健性。
- 建立统一的渐近线性性与时间上的弱收敛性,以支持对生存曲线的推断。
- 在统一框架下处理离散和连续的时间至事件结果。
- 提出一种集成学习策略,用于组合多个条件生存函数的候选估计量,以提升性能。
提出的方法
- 使用交叉拟合以减少基于机器学习的干扰项估计量的过拟合,确保渐近正态性。
- 采用双重稳健估计方程,结合结果回归与逆概率加权,即使仅条件生存或右删失机制被一致估计,估计量仍保持一致。
- 推导出用于通过最小化涉及观测事件指标和逆概率权重的估计方程来估计条件生存函数的损失函数。
- 应用富比尼定理以证明时间积分与条件期望之间的可交换性,确保估计方程的有效性。
- 引入一种集成学习器,用于组合多个条件生存函数的候选估计量,从而增强稳健性和效率。
- 建立最终估计量在时间上渐近线性且一致收敛的条件,支持有效置信带与假设检验。
实验结果
研究问题
- RQ1是否可以使用基于机器学习的条件生存函数估计量来构建具有有效推断能力的处理特定治疗生存曲线的双重稳健估计量?
- RQ2在何种正则性条件下,所提出的估计量在时间上具有统一的渐近线性性?
- RQ3如何利用集成学习将多个生存估计量组合,同时保持双重稳健性与渐近正态性?
- RQ4当事件发生在离散或连续时间时,该方法是否仍保持一致性和有效推断?
- RQ5即使使用灵活的机器学习方法进行干扰项估计,所提出的估计量是否仍能达到参数型收敛速率?
主要发现
- 在较弱的正则性条件下,所提出的交叉拟合双重稳健估计量具有统一的渐近线性性,支持在整个时间区间内有效推断。
- 估计量弱收敛于一个紧致的零均值高斯过程,支持对生存曲线构造置信带。
- 若条件生存函数或右删失机制之一被一致估计,该方法保持一致性,确保对模型误设的稳健性。
- 集成学习器能有效组合多个候选估计量,提升有限样本性能,同时不牺牲理论保证。
- 理论结果适用于离散和连续时间至事件结果,扩大了在真实世界数据中的适用范围。
- 数值研究与对腮腺癌手术的真实世界应用表明,该方法具有实际效用,并优于现有方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。