[Paper Review] Empirical Risk Minimization under Random Censorship: Theory and Practice
This paper proposes a Kaplan-Meier-based plug-in estimator of the empirical risk under random right censorship, enabling empirical risk minimization in regression with censored survival data. Under mild conditions, the learning rate of minimizers of this biased/weighted risk functional achieves the optimal rate of $ O_{\mathbb{P}}(\sqrt{\log n / n}) $, matching the performance of uncensored ERM when model bias from plug-in estimation is ignored.
We consider the classic supervised learning problem, where a continuous non-negative random label $Y$ (i.e. a random duration) is to be predicted based upon observing a random vector $X$ valued in $\mathbb{R}^d$ with $d\geq 1$ by means of a regression rule with minimum least square error. In various applications, ranging from industrial quality control to public health through credit risk analysis for instance, training observations can be right censored, meaning that, rather than on independent copies of $(X,Y)$, statistical learning relies on a collection of $n\geq 1$ independent realizations of the triplet $(X, \; \min\{Y,\; C\},\; δ)$, where $C$ is a nonnegative r.v. with unknown distribution, modeling censorship and $δ=\mathbb{I}\{Y\leq C\}$ indicates whether the duration is right censored or not. As ignoring censorship in the risk computation may clearly lead to a severe underestimation of the target duration and jeopardize prediction, we propose to consider a plug-in estimate of the true risk based on a Kaplan-Meier estimator of the conditional survival function of the censorship $C$ given $X$, referred to as Kaplan-Meier risk, in order to perform empirical risk minimization. It is established, under mild conditions, that the learning rate of minimizers of this biased/weighted empirical risk functional is of order $O_{\mathbb{P}}(\sqrt{\log(n)/n})$ when ignoring model bias issues inherent to plug-in estimation, as can be attained in absence of censorship. Beyond theoretical results, numerical experiments are presented in order to illustrate the relevance of the approach developed.
Motivation & Objective
- To address the challenge of empirical risk minimization in regression when training data are subject to random right censorship, where true labels $ Y $ are unobserved beyond a random censoring time $ C $.
- To develop a theoretically grounded, plug-in-based estimator of the true risk that accounts for censoring via the Kaplan-Meier estimator of the conditional survival function of $ C $ given $ X $.
- To establish nonasymptotic generalization bounds for the resulting weighted empirical risk functional, ensuring learning rates close to optimal even under censorship.
- To validate the method empirically across diverse learning models (linear regression, SVR, random forests) under varying dimensions and censoring rates.
Proposed method
- Proposes a plug-in estimate of the true risk by replacing the unobserved survival function of censorship $ C $ given $ X $ with a nonparametric Kaplan-Meier estimator, resulting in a weighted empirical risk functional.
- Defines the Kaplan-Meier risk as a weighted version of the standard empirical risk, where weights are derived from the inverse of the estimated conditional survival function of $ C $ given $ X $.
- Uses $ U $-process theory and maximal deviation bounds to analyze the empirical process of the weighted risk, establishing concentration inequalities for the risk estimator.
- Applies Corollary 8 on degenerate $ U $-processes to control the deviation of the weighted empirical process, leveraging VC-type complexity and boundedness conditions.
- Derives nonasymptotic bounds on the supremum deviation of the estimated risk from its expectation, using kernel density estimation and uniform entropy conditions.
- Employs a double-robust weighting scheme that accounts for both the censoring indicator $ \delta $ and the conditional survival function, ensuring consistent risk estimation under mild regularity.
Experimental results
Research questions
- RQ1Can empirical risk minimization be consistently extended to regression problems with right-censored response variables, where the true labels are unobserved beyond a random censoring time?
- RQ2What is the optimal learning rate achievable by empirical risk minimizers when using a plug-in estimator of the risk that accounts for censoring via the Kaplan-Meier estimator?
- RQ3How does the performance of the proposed Kaplan-Meier risk estimator compare to naive empirical risk minimization that ignores censoring, especially in high-dimensional input spaces?
- RQ4What nonasymptotic generalization bounds can be established for the weighted empirical risk functional under censorship?
Key findings
- The learning rate of minimizers of the Kaplan-Meier risk functional is $ O_{\mathbb{P}}(\sqrt{\log n / n}) $, matching the optimal rate achievable in the absence of censorship, under mild regularity conditions.
- The proposed method achieves consistent risk estimation by using a plug-in estimator based on the Kaplan-Meier estimator of the conditional survival function of the censoring variable $ C $ given $ X $.
- Numerical experiments show that the IPCW-based estimator significantly outperforms naive ERM methods that ignore censoring, especially under high censoring rates and in higher dimensions.
- The method maintains strong generalization performance across diverse models, including linear regression, support vector regression, and random forests, with consistent improvements in $ L^2 $ error.
- Theoretical analysis confirms that the weighted empirical risk process is well-controlled via $ U $-process theory, with deviation bounds that scale as $ O(\sqrt{\log n / n}) $.
- The approach is robust to model bias from plug-in estimation, as the learning rate remains optimal even when such bias is ignored in the theoretical analysis.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.