Skip to main content
QUICK REVIEW

[论文解读] On Dealing with Censored Largest Observations under Weighted Least Squares

Md Hasinur Rahaman Khan, J.E. Shaw|arXiv (Cornell University)|Dec 9, 2013
Statistical Methods and Inference参考文献 10被引用 4
一句话总结

本文提出五种改进的插补方法,用于加速失效时间(AFT)模型中最大右删失观测值的处理,采用加权最小二乘法,解决了Efron再分配方法带来的偏倚与效率低下问题。这些方法基于Buckley-James插补结合Efron的尾部校正,并引入新颖的均值插补技术,通过二次规划优化的惩罚加权最小二乘法实现,显著降低了均方误差与偏倚,尤其在重度删失情况下表现更优。

ABSTRACT

When observations are subject to right censoring, weighted least squares with appropriate weights (to adjust for censoring) is sometimes used for parameter estimation. With Stute's weighted least squares method, when the largest observation is censored ($Y_{(n)}^+$), it is natural to apply the redistribution to the right algorithm of Efron (1967). However, Efron's redistribution algorithm can lead to bias and inefficiency in estimation. This study explains the issues clearly and proposes some alternative ways of treating $Y_{(n)}^+$. The first four proposed approaches are based on the well known Buckley--James (1979) method of imputation with the Efron's tail correction and the last approach is indirectly based on a general mean imputation technique in literature. All the new schemes use penalized weighted least squares optimized by quadratic programming implemented with the accelerated failure time models. Furthermore, two novel additional imputation approaches are proposed to impute the tail tied censored observations that are often found in survival analysis with heavy censoring. Several simulation studies and real data analysis demonstrated that the proposed approaches generally outperform Efron's redistribution approach and lead to considerably smaller mean squared error and bias estimates.

研究动机与目标

  • 解决在加权最小二乘估计中,处理最大右删失观测值时,Efron的向右再分配算法引入的偏倚与效率低下问题。
  • 开发针对最大右删失观测值的改进插补技术,尤其适用于重度删失和删失数据存在结集的情况。
  • 通过替换Efron的方法为更精确的插补策略,提升加速失效时间(AFT)模型中的参数估计精度。
  • 提供一个公开可用的R包imputeYn,用于在生存分析中实现所提出的方法。
  • 在不同删失水平和协变量相关结构下,评估新插补方法的性能表现。

提出的方法

  • 提出五种基于Buckley-James插补结合Efron尾部校正的插补技术,采用通过二次规划优化的惩罚加权最小二乘法。
  • 引入两种额外的插补方法——迭代法与外推法,用于处理生存数据中常见的重度删失导致的尾部结集删失观测值。
  • 应用Stute的加权最小二乘法(SWLS),采用Kaplan-Meier权重,以考虑删失影响,确保AFT模型估计中的正确加权。
  • 使用条件均值插补、基于重抽样的条件均值插补以及预测差异法对最大右删失观测值$Y_{(n)}^+$进行插补。
  • 在对数变换的生存时间与协变量基础上使用加速失效时间(AFT)模型框架,通过惩罚加权最小二乘法估计参数。
  • 采用二次规划求解惩罚加权最小二乘法中的优化问题,确保数值稳定性和收敛性。

实验结果

研究问题

  • RQ1不同插补策略在最大右删失观测值上的表现,与Efron的再分配方法相比,在偏倚与均方误差方面有何差异?
  • RQ2所提出的插补方法在不同删失水平和协变量相关结构下的性能表现如何?
  • RQ3针对尾部结集删失观测值的新型插补技术,是否能提升重度删失AFT模型中的估计精度?
  • RQ4在真实生存数据(如Channing House数据集)中,所提出方法与Efron方法相比表现如何?
  • RQ5在AFT模型的惩罚加权最小二乘估计中,哪种插补方法能产生最高效且最无偏的参数估计?

主要发现

  • 所提出的插补方法,尤其是条件均值添加法与基于重抽样的条件均值添加法,在所有删失水平下均表现出最低的偏倚与均方误差。
  • 在高删失水平下,预测差异插补法优于Efron的再分配方法;而在低删失与中等删失水平下,两者性能相近。
  • 迭代插补法生成的结集插补值接近预测差异估计值(例如约137.9),表明在结集删失数据中具有稳定性。
  • 外推插补法生成的插补值差异较大(如在Channing House数据中,范围为134.23至200.32),表明可能存在过度离散化,可靠性较低。
  • 在Channing House数据中,年龄的估计系数从迭代法的-0.154变为外推法的-0.218,显示出对模型推断的显著影响。
  • 真实数据分析表明,外推法在生存曲线估计方面优于迭代法,尽管两者在模型拟合度与参数精度方面均优于Efron的方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。