[Paper Review] Statistical Learning of Value-at-Risk and Expected Shortfall
This paper develops nonasymptotic, distribution-free confidence bounds for nonparametric regression under general loss functions using Rademacher complexity and Vapnik-Chervonenkis (VC) theory. It establishes high-probability bounds on estimation error for value-at-risk and expected shortfall approximations, extending prior work to nonstationary and dependent learning samples with explicit concentration inequalities and optimality guarantees for least squares regression.
We propose a non-asymptotic convergence analysis of a two-step approach to learn a conditional value-at-risk (VaR) and a conditional expected shortfall (ES) using Rademacher bounds, in a non-parametric setup allowing for heavy-tails on the financial loss. Our approach for the VaR is extended to the problem of learning at once multiple VaRs corresponding to different quantile levels. This results in efficient learning schemes based on neural network quantile and least-squares regressions. An a posteriori Monte Carlo procedure is introduced to estimate distances to the ground-truth VaR and ES. This is illustrated by numerical experiments in a Student-$t$ toy model and a financial case study where the objective is to learn a dynamic initial margin.
Motivation & Objective
- To derive nonasymptotic, high-probability bounds for nonparametric regression with general loss functions in nonstationary and dependent learning samples.
- To extend Rademacher complexity-based error bounds to dependent data via coupling techniques, improving on i.i.d. assumptions.
- To establish optimality in terms of average $L^2$-distance for least squares regression under nonstationary and dependent sampling.
- To provide a self-contained, accessible reference for Rademacher theory in nonstationary settings, addressing a gap in the literature.
- To support the convergence analysis of value-at-risk and expected shortfall estimation methods proposed in [barcregobngusaa].
Proposed method
- Uses Rademacher complexity theory to derive high-probability upper bounds on the deviation of empirical risk from expected risk for general loss functions.
- Applies concentration inequalities via empirical process theory to control the deviation of regression estimates from conditional expectations.
- Introduces a truncation-based approach to handle heavy-tailed or unbounded responses, enabling bounds under $L^2$-integrability.
- Employs coupling techniques from [BG20] to extend independence-based bounds to $eta$-mixing dependent samples.
- Derives a key inequality (Corollary 4.1) that bounds the average $L^2$-distance between the estimated function and the true conditional expectation, incorporating bias from truncation and variance from sampling.
- Uses the inequality $(a+b)^2 \leq (1+\eta)a^2 + (1+1/\eta)b^2$ to decouple estimation error into bias and variance components for robust bounding.
Experimental results
Research questions
- RQ1Can nonasymptotic confidence bounds be derived for nonparametric regression with general loss functions under nonstationary and dependent training samples?
- RQ2How can Rademacher complexity be extended to dependent data to provide high-probability bounds on estimation error?
- RQ3What is the optimal rate of convergence for least squares regression in nonstationary and dependent settings, and how does it compare to the oracle risk?
- RQ4How can truncation and coupling techniques be combined to control bias and variance in heavy-tailed response settings?
- RQ5What is the role of the Vapnik-Chervonenkis theory in establishing optimality of least squares regression under general sampling assumptions?
Key findings
- The paper establishes a high-probability bound (4.63) on the average $L^2$-distance between the estimated regression function $\hat{g}$ and the true conditional expectation $\Phi_k$, valid with probability at least $1-\delta$.
- The bound (4.63) decomposes into three components: a term proportional to the infimum risk over the function class, a term capturing sampling deviation via $\epsilon_n(c,\lambda)$, and a bias term from truncation involving $\mathbb{E}[(|W_k|-B)^2 \mathbf{1}_{\{|W_k|>B\}}]$.
- The bound is robust to dependence: the coupling method from [BG20] allows extension of independence-based Rademacher bounds to $\beta$-mixing samples.
- The result (4.63) is tight in the sense that it recovers the i.i.d. case as $B \to \infty$ and $\eta, \eta' \to 0$, reducing to the untruncated case.
- The paper shows that the least squares regression scheme achieves optimal $L^2$-risk up to a constant factor, with the bound scaling as $O(\log(1/\delta)/n)$ under mild moment conditions.
- The analysis provides a rigorous foundation for the convergence of value-at-risk and expected shortfall estimators in [barcregobngusaa], particularly for the pinball and least squares schemes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.