Skip to main content
QUICK REVIEW

[论文解读] On redundancy of memoryless sources over countable alphabets

Maryam Hosseini, Narayana Santhanam|arXiv (Cornell University)|Mar 31, 2014
Algorithms and Data Compression参考文献 6被引用 3
一句话总结

本文研究了在可数无限字母表上无 memory 的信源中通用压缩的冗余性,建立了当序列长度增加时,i.i.d. 序列的每符号冗余趋于零的条件。结果表明,有限单字母冗余并不能保证每符号冗余随序列增长而减小,且提出了一个涉及分布尾部行为的充分条件,以确保冗余的次线性增长,从而解决了大字母表估计与压缩理论中的关键空白。

ABSTRACT

The minimum average number of bits need to describe a random variable is its entropy, assuming knowledge of the underlying statistics On the other hand, universal compression supposes that the distribution of the random variable, while unknown, belongs to a known set $\cal P$ of distributions. Such universal descriptions for the random variable are agnostic to the identity of the distribution in $\cal P$. But because they are not matched exactly to the underlying distribution of the random variable, the average number of bits they use is higher, and the excess over the entropy used is the "redundancy". This formulation is fundamental to problems not just in compression, but also estimation and prediction and has a wide variety of applications from language modeling to insurance. In this paper, we study the redundancy of universal encodings of strings generated by independent identically distributed (iid) sampling from a set $\cal P$ of distributions over a countable support. We first show that if describing a single sample from $\cal P$ incurs finite redundancy, then $\cal P$ is tight but that the converse does not always hold. If a single sample can be described with finite worst-case-regret (a more stringent formulation than redundancy above), then it is known that length-$n$ iid samples only incurs a diminishing (in $n$) redundancy per symbol as $n$ increases. However, we show it is possible that a collection $\cal P$ incurs finite redundancy, yet description of length-$n$ iid samples incurs a constant redundancy per symbol encoded. We then show a sufficient condition on $\cal P$ such that length-$n$ iid samples will incur diminishing redundancy per symbol encoded.

研究动机与目标

  • 理解在可数无限字母表上,来自无 memory 信源的 i.i.d. 序列在何种条件下可实现每符号冗余随序列长度增加而趋于零。
  • 澄清有限单字母冗余与长序列下冗余渐近行为之间的关系。
  • 为分布类提供一个充分条件,以确保冗余的次线性增长,从而在高维或无限字母表设置下实现有效的压缩与估计。
  • 通过聚焦于数据导出的一致性与模型特定的冗余衰减速率,解决现有表述(如强冗余与基于模式的压缩)的局限性。

提出的方法

  • 基于尾部概率与熵贡献,将序列分解为‘好’与‘坏’集合,使用基于 log(j)/j 的阈值(j 为序列长度)。
  • 应用一种通用编码方案,分别对‘好’与‘坏’集合上的冗余进行有界处理,利用整个模型类上的统一有界性。
  • 通过构造来自单个分布质量上确界的参考分布 q(n),以有界 KL 散度并确保有限单字母冗余。
  • 使用一个关键不等式,涉及在小概率尾部上对 p(x)log(1/p(x)) 的求和,表明若该和在 δ→0 时趋于零,则冗余减小。
  • 在两种不同条件下分析冗余行为:一种是尾部熵趋于零,另一种是尾部 KL 散度趋于零,证明二者非等价。
  • 构造显式反例(如类 𝒰),表明两种条件互不蕴含,且两者均可与次线性冗余增长共存。

实验结果

研究问题

  • RQ1在可数无限字母表上的分布类满足何种条件时,i.i.d. 序列的每符号冗余随序列长度增加而趋于零?
  • RQ2有限单字母冗余(在类上)是否意味着长序列下每符号冗余趋于零?若否,还需何种额外条件?
  • RQ3是否存在一类分布,其具有有限单字母冗余,却在长 i.i.d. 序列下产生恒定的每符号冗余?若存在,此类行为应如何表征?
  • RQ4是否存在关于分布尾部行为的充分条件,可保证长序列下冗余的次线性增长?该条件与强冗余或基于模式压缩等现有表述有何关联?
  • RQ5关于每符号冗余趋于零的条件是否既必要又充分?当前表征中是否存在空白?

主要发现

  • 有限单字母冗余并不蕴含每符号冗余趋于零;存在具有有限单字母冗余的类,其在长序列下仍保持恒定的每符号冗余。
  • 每符号冗余趋于零的充分条件是:在类上对尾部 T_{p,δ} 上 p(x)log(1/p(x)) 的上确界和在 δ→0 时趋于零。
  • 类 𝒰(其分布 p_k 在 0 处有一个大原子,并在 2^{k²} 个符号上均匀分配质量)具有有限单字母冗余,但不满足尾部熵趋于零的条件。
  • 对于类 𝒰,尾部熵和保持远离零,但长度 n 的冗余仍呈次线性增长,表明该充分条件非必要。
  • 本文构造了一类分布,其尾部熵和远离零,但尾部 KL 散度和趋于零,证明了这两个条件相互独立。
  • 结果表明,现有表述(如强冗余或基于模式的压缩)无法捕捉高维或无限字母表设置下冗余的精细行为,因此需要基于数据导出一致性与模型特定收敛速率的新表征。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。