Skip to main content
QUICK REVIEW

[论文解读] Adaptive estimation of the baseline hazard function in the Cox model by model selection, with high-dimensional covariates

Agathe Guilloux, Sarah Lemler|arXiv (Cornell University)|Mar 1, 2015
Statistical Methods and Inference参考文献 58被引用 5
一句话总结

本文提出了一种在高维Cox比例风险模型中对基线风险函数进行自适应估计的方法,采用两步程序:首先通过部分对数似然函数对回归参数进行Lasso估计,然后通过最小二乘型对比函数进行模型选择以估计基线风险。关键贡献在于建立了一个非渐近的Oracle不等式,该不等式确定了收敛速率,表明随着维度增加,估计偏差也随之增大。

ABSTRACT

The purpose of this article is to provide an adaptive estimator of the baseline function in the Cox model with high-dimensional covariates. We consider a two-step procedure : first, we estimate the regression parameter of the Cox model via a Lasso procedure based on the partial log-likelihood, secondly, we plug this Lasso estimator into a least-squares type criterion and then perform a model selection procedure to obtain an adaptive penalized contrast estimator of the baseline function. Using non-asymptotic estimation results stated for the Lasso estimator of the regression parameter, we establish a non-asymptotic oracle inequality for this penalized contrast estimator of the baseline function, which highlights the discrepancy of the rate of convergence when the dimension of the covariates increases.

研究动机与目标

  • 填补高维生存模型中基线风险函数非渐近估计结果的空白。
  • 在协变量数量p超过样本量n时,提出一种两步程序以估计基线风险函数。
  • 利用模型选择与集中不等式,为基线风险函数估计器提供理论保证。
  • 建立一个非渐近Oracle不等式,量化在高维稀疏条件下估计器的收敛速率。
  • 通过结合Lasso进行回归参数估计与惩罚型模型选择,确保基线风险函数估计器的自适应性。

提出的方法

  • 在部分对数似然函数上使用Lasso惩罚来估计高维回归参数β₀。
  • 将β₀的Lasso估计量代入最小二乘型对比函数,以估计基线风险。
  • 基于AIC与Mallows准则的模型选择程序,从一组候选函数中选择最优的基线风险估计器。
  • 利用非渐近集中不等式及Huang等(2013)的结果,控制Lasso估计量的偏差。
  • 推导基线风险函数惩罚对比估计器的非渐近Oracle不等式。
  • 利用Doob-Meyer分解与鞅的Bernstein不等式,界定估计误差中经验过程项的上界。

实验结果

研究问题

  • RQ1当协变量数量p超过样本量n时,如何在Cox模型中一致地估计基线风险函数?
  • RQ2在高维设置下,基线风险函数自适应估计器的非渐近收敛速率是什么?
  • RQ3对最小二乘对比函数应用模型选择程序,能否为基线风险估计获得Oracle不等式?
  • RQ4真实回归参数β₀的稀疏性如何影响基线风险函数估计的准确性?
  • RQ5对于Lasso估计后接模型选择的联合两步程序,可建立哪些理论保证?

主要发现

  • 所提出的估计器实现了非渐近Oracle不等式,有效控制了基线风险函数的估计误差。
  • 估计器的收敛速率随着维度增加而恶化,这体现在误差界中对log(pnᵏ)/n的依赖。
  • 以超过1 - cn⁻ᵏ的概率,β₀的Lasso估计量满足|β̂ - β₀|₁ ≤ C(s)√(log(pnᵏ)/n),其中s为稀疏性指标。
  • 在设计矩阵与基线风险函数有界性的假设下,推导出基线风险估计器的Oracle不等式。
  • 通过从候选函数类中选择最优模型,该方法确保了自适应性,当真实模型位于候选集中时,可达到最优估计速率。
  • 理论结果为非渐近性质,即使在p > n时依然成立,将现有方法扩展至经典低维范式之外。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。