Skip to main content
QUICK REVIEW

[论文解读] Stable variable selection for right censored data: comparison of methods

Marie Walschaerts, Eve Leconte|arXiv (Cornell University)|Mar 22, 2012
Statistical Methods and Inference参考文献 40被引用 8
一句话总结

本文提出并比较了针对高维右删失生存数据中变量选择的基于自展法的稳定化方法,重点研究Cox比例风险模型与生存树。研究发现,Cox模型中的自展Lasso方法以及生存树中的节点级稳定化方法在预测准确率与临床可解释性之间取得了良好平衡,其可解释性优于随机生存森林,同时在低样本量情况下仍保持具有竞争力的性能。

ABSTRACT

The instability in the selection of models is a major concern with data sets containing a large number of covariates. This paper deals with variable selection methodology in the case of high-dimensional problems where the response variable can be right censored. We focuse on new stable variable selection methods based on bootstrap for two methodologies: the Cox proportional hazard model and survival trees. As far as the Cox model is concerned, we investigate the bootstrapping applied to two variable selection techniques: the stepwise algorithm based on the AIC criterion and the L1-penalization of Lasso. Regarding survival trees, we review two methodologies: the bootstrap node-level stabilization and random survival forests. We apply these different approaches to two real data sets. We compare the methods on the prediction error rate based on the Harrell concordance index and the relevance of the interpretation of the corresponding selected models. The aim is to find a compromise between a good prediction performance and ease to interpretation for clinicians. Results suggest that in the case of a small number of individuals, a bootstrapping adapted to L1-penalization in the Cox model or a bootstrap node-level stabilization in survival trees give a good alternative to the random survival forest methodology, known to give the smallest prediction error rate but difficult to interprete by non-statisticians. In a clinical perspective, the complementarity between the methods based on the Cox model and those based on survival trees would permit to built reliable models easy to interprete by the clinician.

研究动机与目标

  • 解决高维右删失生存数据中变量选择的不稳定性问题。
  • 评估基于自展法的稳定化方法在Cox模型与生存树中的应用,以提升模型的可靠性与可解释性。
  • 利用真实乳腺癌与不孕症数据集,比较不同方法在预测性能与临床可解释性方面的表现。
  • 为临床医生识别预测准确率与可解释性之间实用的折中方案。

提出的方法

  • 在Cox模型中对Lasso惩罚进行自展法处理,以稳定变量选择,称为BLS(Bootstrap Lasso Selection)。
  • 在生存树中采用节点级自展法稳定化(BSS),以减少过拟合并提升选择稳定性。
  • 将随机生存森林(RSF)作为预测性能的基准。
  • 采用逐步AIC选择法与单棵生存树作为基线方法。
  • 使用Harrell的 concordance index 评估不同方法的预测误差率。
  • 在两个真实数据集上验证结果:一个经典的乳腺癌基因表达数据集与一个具有复杂协变量的原创不孕症数据集。

实验结果

研究问题

  • RQ1自展法如何提升高维右删失生存数据中变量选择的稳定性?
  • RQ2在Cox模型方法与树模型方法之间,哪种方法在预测准确率与临床可解释性之间提供了最佳平衡?
  • RQ3在生存树中采用节点级自展法稳定化是否能在不牺牲性能的前提下,优于标准生存树与随机生存森林,从而提升可解释性?
  • RQ4在事件数相对于协变量较少的数据集中,不同稳定化技术的表现如何?
  • RQ5基于自展法的方法能否可靠识别出生殖医学中具有临床意义的因素,如女性年龄与不孕持续时间?

主要发现

  • BLS方法(Cox模型中的自展Lasso)在预测误差率上与随机生存森林相当,同时生成单一、可解释的模型,因此特别适合临床应用。
  • 在乳腺癌数据集中,生存树的节点级自展法稳定化(BSS)未提升性能,可能因事件数过少;但在样本量充足的不孕症数据集中展现出潜力。
  • BNLS(Bootstrap Node-Level Stabilization)方法在任一数据集中均未优于单棵生存树,可能由于复杂度参数难以调优。
  • 所有方法在不孕症数据集中均识别出女性年龄 >34.5岁与不孕持续时间 >33.5个月为关键预测因子,证实其临床相关性。
  • 男性因素如精索静脉曲张与睾丸体积在BLS、BSS与BNLS中均被一致选中,凸显其生物学相关性,而这些因素在以往研究中常被忽略。
  • 研究结论认为,应互补使用基于Cox模型与基于树的方法,以充分发挥其在预测性能与可解释性方面的各自优势。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。