Skip to main content
QUICK REVIEW

[Paper Review] On redundancy of memoryless sources over countable alphabets

Maryam Hosseini, Narayana Santhanam|arXiv (Cornell University)|Mar 31, 2014
Algorithms and Data Compression6 references3 citations
TL;DR

This paper investigates the redundancy of universal compression for memoryless sources over countably infinite alphabets, establishing conditions under which the per-symbol redundancy of i.i.d. sequences diminishes to zero as sequence length increases. It shows that finite single-letter redundancy does not guarantee diminishing per-symbol redundancy, and provides a sufficient condition involving tail behavior of distributions to ensure sublinear growth of redundancy, resolving key gaps in large-alphabet estimation and compression theory.

ABSTRACT

The minimum average number of bits need to describe a random variable is its entropy, assuming knowledge of the underlying statistics On the other hand, universal compression supposes that the distribution of the random variable, while unknown, belongs to a known set $\cal P$ of distributions. Such universal descriptions for the random variable are agnostic to the identity of the distribution in $\cal P$. But because they are not matched exactly to the underlying distribution of the random variable, the average number of bits they use is higher, and the excess over the entropy used is the "redundancy". This formulation is fundamental to problems not just in compression, but also estimation and prediction and has a wide variety of applications from language modeling to insurance. In this paper, we study the redundancy of universal encodings of strings generated by independent identically distributed (iid) sampling from a set $\cal P$ of distributions over a countable support. We first show that if describing a single sample from $\cal P$ incurs finite redundancy, then $\cal P$ is tight but that the converse does not always hold. If a single sample can be described with finite worst-case-regret (a more stringent formulation than redundancy above), then it is known that length-$n$ iid samples only incurs a diminishing (in $n$) redundancy per symbol as $n$ increases. However, we show it is possible that a collection $\cal P$ incurs finite redundancy, yet description of length-$n$ iid samples incurs a constant redundancy per symbol encoded. We then show a sufficient condition on $\cal P$ such that length-$n$ iid samples will incur diminishing redundancy per symbol encoded.

Motivation & Objective

  • To understand the conditions under which universal compression of i.i.d. sequences from memoryless sources over countably infinite alphabets achieves diminishing per-symbol redundancy.
  • To clarify the relationship between finite single-letter redundancy and the asymptotic behavior of redundancy over long sequences.
  • To provide a sufficient condition on the distribution class that ensures sublinear growth of redundancy, enabling effective compression and estimation in high- or infinite-alphabet settings.
  • To address limitations in existing formulations—such as strong redundancy and pattern-based compression—by focusing on data-derived consistency and model-specific redundancy decay rates.

Proposed method

  • Introduces a decomposition of sequences into 'good' and 'bad' sets based on tail probabilities and entropy contributions, using thresholds derived from log(j)/j for sequence length j.
  • Applies a universal coding scheme that separately bounds redundancy over the 'good' and 'bad' sets, leveraging uniform bounds across the entire model class.
  • Uses a reference distribution q(n) constructed from suprema of individual distribution masses to bound KL divergence and ensure finite single-letter redundancy.
  • Employs a key inequality involving the sum of p(x)log(1/p(x)) over small-probability tails, showing that if this sum vanishes as δ→0, redundancy diminishes.
  • Analyzes the behavior of redundancy under two distinct conditions: one where tail entropy vanishes and another where tail KL divergence vanishes, demonstrating their non-equivalence.
  • Constructs explicit counterexamples (e.g., class 𝒰) to show that neither condition implies the other, and that both can coexist with sublinear redundancy growth.

Experimental results

Research questions

  • RQ1Under what conditions on a class of distributions over a countably infinite alphabet does the per-symbol redundancy of i.i.d. sequences diminish to zero as sequence length increases?
  • RQ2Does finite single-letter redundancy (over the class) imply diminishing per-symbol redundancy for long sequences, and if not, what additional conditions are required?
  • RQ3Can a class of distributions have finite single-letter redundancy yet incur constant per-symbol redundancy for long i.i.d. sequences, and if so, how can such behavior be characterized?
  • RQ4Is there a sufficient condition on the tail behavior of distributions that guarantees sublinear growth of redundancy over long sequences, and how does it relate to existing formulations like strong redundancy or pattern-based compression?
  • RQ5Are the conditions for diminishing per-symbol redundancy both necessary and sufficient, or are there gaps in current characterizations?

Key findings

  • Finite single-letter redundancy does not imply diminishing per-symbol redundancy; there exist classes with finite single-letter redundancy that still incur constant per-symbol redundancy for long sequences.
  • A sufficient condition for diminishing per-symbol redundancy is that the supremum over the class of the sum of p(x)log(1/p(x)) over tails T_{p,δ} tends to zero as δ→0.
  • The class 𝒰, constructed with distributions p_k having a large atom at 0 and uniform mass over 2^{k²} symbols, has finite single-letter redundancy but fails the tail-entropy vanishing condition.
  • For the class 𝒰, the tail-entropy sum remains bounded away from zero, yet the length-n redundancy grows sublinearly, showing that the sufficient condition is not necessary.
  • The paper constructs a class where the tail-entropy sum is bounded away from zero but the tail KL divergence sum vanishes, demonstrating that the two conditions are independent.
  • The results imply that existing formulations—such as strong redundancy or pattern-based compression—fail to capture the nuanced behavior of redundancy in high- or infinite-alphabet settings, necessitating new characterizations based on data-derived consistency and model-specific convergence rates.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.