Skip to main content
QUICK REVIEW

[Paper Review] On Dealing with Censored Largest Observations under Weighted Least Squares

Md Hasinur Rahaman Khan, J.E. Shaw|arXiv (Cornell University)|Dec 9, 2013
Statistical Methods and Inference10 references4 citations
TL;DR

This paper proposes five improved imputation methods for the largest censored observations in accelerated failure time (AFT) models using weighted least squares, addressing bias and inefficiency from Efron’s redistribution method. The approaches—based on Buckley-James imputation with Efron’s tail correction and novel mean imputation techniques—use penalized weighted least squares via quadratic programming and outperform Efron’s method with significantly reduced mean squared error and bias, especially under heavy censoring.

ABSTRACT

When observations are subject to right censoring, weighted least squares with appropriate weights (to adjust for censoring) is sometimes used for parameter estimation. With Stute's weighted least squares method, when the largest observation is censored ($Y_{(n)}^+$), it is natural to apply the redistribution to the right algorithm of Efron (1967). However, Efron's redistribution algorithm can lead to bias and inefficiency in estimation. This study explains the issues clearly and proposes some alternative ways of treating $Y_{(n)}^+$. The first four proposed approaches are based on the well known Buckley--James (1979) method of imputation with the Efron's tail correction and the last approach is indirectly based on a general mean imputation technique in literature. All the new schemes use penalized weighted least squares optimized by quadratic programming implemented with the accelerated failure time models. Furthermore, two novel additional imputation approaches are proposed to impute the tail tied censored observations that are often found in survival analysis with heavy censoring. Several simulation studies and real data analysis demonstrated that the proposed approaches generally outperform Efron's redistribution approach and lead to considerably smaller mean squared error and bias estimates.

Motivation & Objective

  • To address the bias and inefficiency introduced by Efron’s redistribution to the right algorithm when handling the largest censored observation in weighted least squares estimation.
  • To develop improved imputation techniques for the largest censored observations, particularly in cases of heavy censoring and tied censored data.
  • To enhance parameter estimation in accelerated failure time (AFT) models by replacing Efron’s method with more accurate imputation strategies.
  • To provide a publicly available R package, imputeYn, for implementing the proposed methods in survival analysis.
  • To evaluate the performance of new imputation methods under varying censoring levels and correlation structures in covariates.

Proposed method

  • Proposes five imputation techniques based on Buckley-James imputation with Efron’s tail correction, using penalized weighted least squares optimized via quadratic programming.
  • Introduces two additional imputation methods—iterative and extrapolation—for handling tail-tied censored observations common in survival data with heavy censoring.
  • Applies Stute’s weighted least squares (SWLS) with Kaplan-Meier weights to account for censoring, ensuring proper weighting in AFT model estimation.
  • Imputes the largest censored observation $Y_{(n)}^+$ using conditional mean imputation, resampling-based conditional mean, and predicted difference approaches.
  • Uses the accelerated failure time (AFT) model framework with log-transformed survival times and covariates, estimating parameters via penalized weighted least squares.
  • Employs quadratic programming to solve the optimization problem in penalized weighted least squares, ensuring numerical stability and convergence.

Experimental results

Research questions

  • RQ1How do different imputation strategies for the largest censored observation compare to Efron’s redistribution method in terms of bias and mean squared error?
  • RQ2What is the performance of the proposed imputation methods under varying levels of censoring and correlation structures in covariates?
  • RQ3Can novel imputation techniques for tail-tied censored observations improve estimation accuracy in AFT models with heavy censoring?
  • RQ4How do the proposed methods compare to Efron’s approach in real-world survival data, such as the Channing House dataset?
  • RQ5Which imputation method yields the most efficient and least biased parameter estimates in penalized weighted least squares estimation under AFT models?

Key findings

  • The proposed imputation methods, especially conditional mean adding and resampling-based conditional mean adding, consistently yield the lowest bias and mean squared error across all censoring levels.
  • Under high censoring, the predicted difference imputation approach outperforms Efron’s redistribution method, while at lower and medium censoring levels, performance is comparable.
  • The iterative imputation method produces tied imputed values close to the predicted difference estimate (e.g., ~137.9), suggesting stability under tied censored data.
  • The extrapolation imputation method generates widely varying imputed values (e.g., 134.23 to 200.32 in the Channing House data), indicating potential overdispersion and less reliability.
  • In the Channing House data, the estimated coefficient for age changed from -0.154 (iterative) to -0.218 (extrapolation), showing substantial impact on model inference.
  • The real data analysis demonstrates that the extrapolation method outperforms the iterative method in survival curve estimation, though both improve upon Efron’s approach in terms of model fit and parameter precision.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.