[Paper Review] Non-asymptotic Oracle Inequalities for the High-Dimensional Cox Regression via Lasso
This paper establishes non-asymptotic oracle inequalities for high-dimensional Cox regression using the lasso penalty, overcoming the challenges of non-i.i.d. and non-Lipschitz partial likelihood loss by approximating the negative log-partial likelihood with an i.i.d. surrogate loss. The key contribution is a finite-sample bound showing that the lasso estimator achieves estimation and prediction error rates comparable to an oracle that knows the true sparse model.
We consider the finite sample properties of the regularized high-dimensional Cox regression via lasso. Existing literature focuses on linear models or generalized linear models with Lipschitz loss functions, where the empirical risk functions are the summations of independent and identically distributed (iid) losses. The summands in the negative log partial likelihood function for censored survival data, however, are neither iid nor Lipschitz. We first approximate the negative log partial likelihood function by a sum of iid non-Lipschitz terms, then derive the non-asymptotic oracle inequalities for the lasso penalized Cox regression using pointwise arguments to tackle the difficulty caused by the lack of iid and Lipschitz property.
Motivation & Objective
- Address the lack of finite-sample theory for high-dimensional Cox regression with lasso regularization, particularly when the loss function lacks i.i.d. and Lipschitz properties.
- Overcome theoretical challenges arising from the non-i.i.d. and non-Lipschitz nature of the negative log-partial likelihood in censored survival data.
- Derive non-asymptotic oracle inequalities that bound the excess risk and estimation error of the lasso estimator in a high-dimensional setting.
- Provide a rigorous finite-sample analysis of variable selection and estimation consistency for the lasso in the Cox model, extending results from generalized linear models to survival analysis.
Proposed method
- Approximate the non-i.i.d. negative log-partial likelihood with a surrogate empirical loss function based on population expectations, creating an i.i.d. structure for theoretical analysis.
- Introduce a working loss function $ \gamma_{f_{\theta}} $ with expected value $ l(\theta) = P\gamma_{f_{\theta}} $, enabling the use of pointwise arguments to handle non-Lipschitz behavior.
- Use a boundedness assumption in place of the Lipschitz condition from van de Geer (2008), allowing analysis under weaker regularity conditions.
- Apply a chaining argument with a geometric progression of balls to control the entropy of the parameter space, replacing the need for Lipschitz continuity.
- Derive high-probability bounds on the excess risk $ \hat{\mathcal{E}} $ and the $ \ell_1 $-norm of the estimation error $ I(\hat{\theta}_n - \theta_n^*) $ using a peeling device and concentration inequalities.
- Handle random weights in the lasso penalty by adapting tail probability bounds from van de Geer (2008), ensuring the final oracle inequality holds with high probability.
Experimental results
Research questions
- RQ1Can non-asymptotic oracle inequalities be established for the lasso in high-dimensional Cox regression despite the non-i.i.d. and non-Lipschitz nature of the partial likelihood?
- RQ2How can the theoretical framework for generalized linear models be adapted to survival models with censored data?
- RQ3What is the finite-sample performance of the lasso estimator in terms of excess risk and estimation error under high-dimensional sparsity?
- RQ4How do random weights in the lasso penalty affect the concentration bounds and the final oracle inequality?
- RQ5Can a surrogate loss function be constructed to preserve the i.i.d. structure while approximating the true negative log-partial likelihood?
Key findings
- The paper establishes a non-asymptotic oracle inequality showing that the excess risk of the lasso estimator satisfies $ \hat{\mathcal{E}} \leq \frac{1}{1-\delta} \epsilon_n^* $ with high probability, where $ \epsilon_n^* $ is a term related to the model complexity and sparsity.
- The estimation error in $ \ell_1 $-norm is bounded by $ I(\hat{\theta}_n - \theta_n^*) \leq d(\delta_1, \delta_2) \frac{\zeta_n^*}{b} $ with high probability, controlling the deviation from the true sparse model.
- The probability of the oracle inequality holding is bounded below by $ 1 - \log_{1+b}\left\{ \frac{(1+b)^2}{\delta_1\delta_2} \cdot \frac{d(\delta_1,\delta_2)(1-\delta^2)}{\delta b} \right\} \times \left\{ (1 + \frac{3}{10}W^2)\exp(-n\bar{a}_n^2 r_1^2) + 2\exp(-n\pi^2/2) \right\} $, which decays exponentially in sample size.
- The results hold under a boundedness assumption replacing the Lipschitz condition, making the framework applicable to the non-Lipschitz partial likelihood of the Cox model.
- The analysis extends to random weights in the lasso penalty by incorporating additional tail probability bounds, ensuring robustness to estimation of the design covariance.
- The derived bounds match the form of oracle inequalities in van de Geer (2008) for generalized linear models, but are adapted to the survival analysis setting with censored data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.