Skip to main content
QUICK REVIEW

[Paper Review] A constrained risk inequality for general losses

John C. Duchi, Feng Ruan|arXiv (Cornell University)|Apr 22, 2018
Statistical Methods and Inference11 references3 citations
TL;DR

This paper presents a general constrained risk inequality for statistical estimation under arbitrary non-decreasing loss functions, extending Brown and Low's result beyond squared error loss. By leveraging the Cauchy-Schwarz inequality and a two-point testing framework, it establishes a lower bound on estimation risk under one distribution given an upper bound under another, enabling finite-sample lower bounds for super-efficient and adaptive estimators across diverse loss types, including absolute error and deviation probabilities.

ABSTRACT

We provide a general constrained risk inequality that applies to arbitrary non-decreasing losses, extending a result of Brown and Low [Ann. Stat. 1996]. Given two distributions $P_0$ and $P_1$, we find a lower bound for the risk of estimating a parameter $θ(P_1)$ under $P_1$ given an upper bound on the risk of estimating the parameter $θ(P_0)$ under $P_0$. The inequality is a useful pedagogical tool, as its proof relies only on the Cauchy-Schwartz inequality, it applies to general losses, and it transparently gives risk lower bounds on super-efficient and adaptive estimators.

Motivation & Objective

  • To develop a general lower bound for estimation risk under arbitrary non-decreasing loss functions, moving beyond the squared error loss used in prior work.
  • To provide a transparent, pedagogically useful proof relying only on the Cauchy-Schwarz inequality, applicable to a broad class of statistical estimation problems.
  • To quantify the trade-off in risk between distributions where an estimator is super-efficient and where it incurs inflated error, illustrating the impossibility of uniform super-efficiency.
  • To extend the applicability of constrained risk inequalities to nonparametric estimation problems and functional estimation, particularly in settings with general loss functions.
  • To offer a unified framework for deriving finite-sample lower bounds in adaptive estimation, bridging the gap toward local asymptotic minimax theory.

Proposed method

  • Formalizes estimation risk using a general loss function $ \ell(\|\widehat{\theta} - \theta(P)\|_2) $, where $ \ell $ is non-decreasing and convex.
  • Introduces the $ \chi^2 $-affinity $ \rho(P_1 \| P_0) = \mathbb{E}_0[(dP_1/dP_0)^2] $ to measure the similarity between two distributions $ P_0 $ and $ P_1 $.
  • Derives a lower bound on the risk $ R(\widehat{\theta}, P_1) $ given an upper bound $ \delta $ on $ R(\widehat{\theta}, P_0) $, using the Cauchy-Schwarz inequality on transformed loss terms.
  • Defines separation $ \Delta = 2\ell(\frac{1}{2}\|\theta_0 - \theta_1\|_2) $ to quantify the distance between parameters under $ P_0 $ and $ P_1 $.
  • Applies a likelihood ratio change of measure to relate expectations under $ P_1 $ to those under $ P_0 $, enabling the derivation of the final risk inequality.
  • Generalizes the result to non-convex losses and different-dimensional parameter spaces via majorization and Hölder’s inequality, yielding Corollaries 1 and 2.

Experimental results

Research questions

  • RQ1Can a constrained risk inequality be established for general non-decreasing loss functions beyond squared error?
  • RQ2What is the minimal risk achievable for estimating $ \theta(P_1) $ when the risk at $ P_0 $ is bounded, under arbitrary loss?
  • RQ3How does the risk of a super-efficient estimator at $ P_0 $ constrain its performance at a nearby distribution $ P_1 $?
  • RQ4To what extent can this inequality be applied to nonparametric estimation problems with general loss functions?
  • RQ5What are the finite-sample implications of this inequality for adaptive and efficient estimators in functional estimation?

Key findings

  • The paper establishes the inequality $ R(\widehat{\theta}, P_1) \geq \left[ \Delta^{1/2} - (\rho(P_1 \| P_0) \cdot \delta)^{1/2} \right]_+^2 $, where $ \Delta = 2\ell(\frac{1}{2}\|\theta_0 - \theta_1\|_2) $, providing a lower bound on risk under $ P_1 $ given a bound $ \delta $ under $ P_0 $.
  • The proof relies solely on the Cauchy-Schwarz inequality, making it transparent and widely applicable across different loss functions.
  • For non-convex losses, Corollary 1 shows the same bound holds with $ \Delta = \ell(\frac{1}{2}\|\theta_0 - \theta_1\|_2) $, extending the result beyond convex losses.
  • In the case of $ k $-dimensional parameters with $ k \leq 2 $, the bound becomes $ R(\widehat{\theta}, P_1) \geq \left[ \Delta^{1/2} - (\rho(P_1 \| P_0) \cdot \delta)^{1/2} \right]_+^2 $, with $ \Delta = \|\theta_0 - \theta_1\|_2 $.
  • For $ k > 2 $, the bound is $ R(\widehat{\theta}, P_1) \geq \left[ \Delta - (\rho(P_1 \| P_0) \cdot \delta^2)^{1/2} \right]_+^k $, derived via Hölder’s inequality and reduction to the $ k=2 $ case.
  • The method is applied to normal mean estimation and nonparametric function estimation, demonstrating its utility in deriving finite-sample lower bounds for adaptive estimators.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.