Skip to main content
QUICK REVIEW

[Paper Review] Novel Bernstein-like Concentration Inequalities for the Missing Mass

Bahman Yari Saeed Khanloo, Gholamreza Haffari|arXiv (Cornell University)|Mar 10, 2015
Statistical Mechanics and Entropy8 references3 citations
TL;DR

This paper introduces novel Bernstein-like concentration inequalities for the missing mass in discrete distributions, using a thresholding technique to control heterogeneity in outcome probabilities. It achieves tighter bounds than prior work—especially for small deviations—by regulating term magnitudes via a novel variance-aware regularization, improving state-of-the-art results for $\epsilon < 0.021$ (lower tail) and $\epsilon < 0.045$ (upper tail).

ABSTRACT

We are concerned with obtaining novel concentration inequalities for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We not only derive - for the first time - distribution-free Bernstein-like deviation bounds with sublinear exponents in deviation size for missing mass, but also improve the results of McAllester and Ortiz (2003) andBerend and Kontorovich (2013, 2012) for small deviations which is the most interesting case in learning theory. It is known that the majority of standard inequalities cannot be directly used to analyze heterogeneous sums i.e. sums whose terms have large difference in magnitude. Our generic and intuitive approach shows that the heterogeneity issue introduced in McAllester and Ortiz (2003) is resolvable at least in the case of missing mass via regulating the terms using our novel thresholding technique.

Motivation & Objective

  • Address the lack of sharp concentration inequalities for the missing mass, particularly in the small-deviation regime critical to learning theory.
  • Overcome the limitations of standard inequalities in handling heterogeneous sums where terms vary widely in magnitude.
  • Develop a distribution-free framework that enables tighter bounds on missing mass fluctuations without relying on specific distributional assumptions.
  • Refine existing bounds from McAllester & Ortiz (2003) and Berend & Kontorovich (2013) by introducing a thresholding mechanism to regulate term contributions.
  • Establish tighter upper and lower deviation bounds for the missing mass using Bernstein-type inequalities with sublinear exponents in deviation size.

Proposed method

  • Introduce a thresholding technique to partition outcomes into low- and high-probability groups based on a threshold $\tau^*$, regulating term magnitudes in the sum.
  • Define truncated random variables $Y'$ and $Y''$ to control variance and martingale differences, enabling application of standard concentration inequalities.
  • Apply Bernstein’s inequality to the truncated sum by bounding the variance $V_{\mathcal{L}''}$ and the maximum term size $\alpha_l$, using $\alpha_l = \tau'$ and $V_{\mathcal{L}''} \leq \frac{\theta}{n} \epsilon + \frac{2\theta}{3n} \cdot \frac{\gamma-1}{\gamma} \epsilon$.
  • Use exponential moment methods with a carefully chosen $\gamma$ to optimize the bound, leading to $\exp\left(-\frac{3n\epsilon(\gamma-1)^2}{10\gamma^2 \ln(\gamma/\epsilon)}\right)$, which is minimized over $\gamma$.
  • Control the compensation gap $g_l(\theta)$ via $|g_l(\theta)| \leq f(\theta)$, ensuring it remains negligible for small $\epsilon$, thus preserving tightness.
  • Establish both upper and lower tail bounds by symmetric application of the method, with the final bound expressed as $\mathbb{P}(Y - \mathbb{E}[Y] \leq -\epsilon) \leq \exp(-c(\epsilon) \cdot n\epsilon)$.

Experimental results

Research questions

  • RQ1Can we derive tighter concentration inequalities for the missing mass that are distribution-free and effective even for small deviations?
  • RQ2How can we resolve the heterogeneity problem in sums of Bernoulli variables with vastly differing probabilities, as seen in missing mass estimation?
  • RQ3To what extent can thresholding techniques improve the sharpness of Bernstein-type bounds in the context of missing mass?
  • RQ4Can we achieve sublinear exponents in $\epsilon$ for the deviation rate while maintaining distribution-free guarantees?
  • RQ5How does the compensation gap introduced by truncation affect the final bound, and can it be made negligible for small $\epsilon$?

Key findings

  • The proposed method achieves tighter lower-tail bounds than Berend and Kontorovich (2013) for all $\epsilon < 0.021$, with the improvement being most pronounced in the small-deviation regime.
  • For the upper tail, the new bounds improve upon the state-of-the-art for all $\epsilon < 0.045$, demonstrating enhanced performance in the most relevant range for learning theory.
  • The compensation gap $|g(\epsilon)|$ is shown to be bounded by $\sqrt{e} \cdot \exp\left(W_{-1}\left(\frac{-\epsilon}{2\sqrt{e}}\right)\right)$, which is negligible for small $\epsilon$, justifying the tightness of the final bounds.
  • The final bound is expressed as $\mathbb{P}(Y - \mathbb{E}[Y] \leq -\epsilon) \leq \exp(-c(\epsilon) \cdot n\epsilon)$, with $c(\epsilon)$ depending on $\epsilon$ and achieving sublinear dependence on $\epsilon$ in the exponent.
  • The method enables direct application of standard concentration inequalities (e.g., Bernstein) to the missing mass by first regulating term heterogeneity via thresholding.
  • The bounds are distribution-free and hold under the mild condition $n \geq \lceil \gamma_\epsilon \rceil - 1$, ensuring applicability in non-asymptotic settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.