[论文解读] Consistency of survival tree and forest models: splitting bias and correction
本文研究了右删失数据下生存树与随机森林模型的一致性,识别出由失效分布与删失分布之间的依赖关系引发的分裂偏差。提出了一种偏差校正的累积风险率分裂规则,其一致性收敛速率仅依赖于失效变量数量,显著优于标准方法的预测误差。
Random survival forest and survival trees are popular models in statistics and machine learning. However, there is a lack of general understanding regarding consistency, splitting rules and influence of the censoring mechanism. In this paper, we investigate the statistical properties of existing methods from several interesting perspectives. First, we show that traditional splitting rules with censored outcomes rely on a biased estimation of the within-node failure distribution. To exactly quantify this bias, we develop a concentration bound of the within-node estimation based on non i.i.d. samples and apply it to the entire forest. Second, we analyze the entanglement between the failure and censoring distributions caused by univariate splits, and show that without correcting the bias at an internal node, survival tree and forest models can still enjoy consistency under suitable conditions. In particular, we demonstrate this property under two cases: a finite-dimensional case where the splitting variables and cutting points are chosen randomly, and a high-dimensional case where the covariates are weakly correlated. Our results can also degenerate into an independent covariate setting, which is commonly used in the random forest literature for high-dimensional sparse models. However, it may not be avoidable that the convergence rate depends on the total number of variables in the failure and censoring distributions. Third, we propose a new splitting rule that compares bias-corrected cumulative hazard functions at each internal node. We show that the rate of consistency of this new model depends only on the number of failure variables, which improves from non-bias-corrected versions. We perform simulation studies to confirm that this can substantially benefit the prediction error.
研究动机与目标
- 理解在失效与删失分布依赖时,生存树与森林模型在右删失结果下的的一致性。
- 识别并量化传统分裂规则因依赖于有偏的节点内失效分布估计而引入的分裂偏差。
- 在有限维与高维设置下,建立生存树与森林在存在该偏差时仍保持一致的条件。
- 提出一种通过比较每个节点处偏差校正的累积风险率函数来校正偏差的新分裂规则。
- 证明新方法在仅依赖失效变量数量的收敛速率下实现一致性,而非总协变量数量。
提出的方法
- 推导了在非独立同分布样本下,节点内累积风险率估计的集中性界,考虑了删失与失效依赖关系。
- 分析了单变量分裂导致的失效与删失分布之间的纠缠,表明即使在一致性成立时偏差仍可能持续存在。
- 提出一种基于偏差校正累积风险率函数 $ \widetilde{\Lambda}_{\mathcal{A}}^{\ast}(t) $ 的新分裂规则,以消除删失分布的影响。
- 在两种情形下建立了新模型的一致性:有限维下采用随机分裂,以及高维下协变量弱相关。
- 利用自适应集中不等式与渐近界,证明新方法的收敛速率仅依赖于失效变量数量,而非总协变量数量。
- 通过模拟研究验证,偏差校正分裂规则相比标准方法显著降低了预测误差。
实验结果
研究问题
- RQ1基于节点内失效分布估计的传统生存树分裂方法,在失效与删失分布依赖时是否能实现一致模型?
- RQ2当分裂规则因删失依赖而产生偏差时,生存树与森林是否仍能在未显式校正的情况下保持一致?
- RQ3协变量总数(尤其是删失相关变量)对生存森林模型收敛速率有何影响?
- RQ4能否构建一种偏差校正分裂规则,以隔离失效信号并消除删失影响,从而提升一致性?
- RQ5所提方法是否比非偏差校正版本具有更快的收敛速率,且是否转化为更低的预测误差?
主要发现
- 传统生存树与森林中的分裂规则在失效与删失分布依赖时,对节点内失效分布的估计存在系统性偏差。
- 在适当条件下,如有限维下的随机分裂或高维下的弱相关协变量,生存树与森林即使在未校正偏差下仍可实现一致性。
- 标准模型的收敛速率依赖于失效与删失机制中涉及的总变量数,可能并非最优。
- 所提出的基于 $ \widetilde{\Lambda}_{\mathcal{A}}^{\ast}(t) $ 的偏差校正分裂规则,实现了一致性,且收敛速率仅依赖于失效变量数量,而非总协变量数。
- 模拟研究证实,偏差校正方法相比标准生存森林模型显著降低了预测误差。
- 理论分析表明,新方法的一致性对删失依赖具有鲁棒性,且在有限维与高维设置下均实现了改进的收敛速率。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。