[论文解读] On principles of large deviation and selected data compression
本文通过结合大偏差原理的效用驱动数据压缩,扩展了香农无噪声编码定理。提出了一种概率框架,通过基于权重函数选择字符串来优化存储,通过在约束条件下最小化熵来推导渐近存储速率,关键结果是渐近体积增长速率由涉及熵和效用约束的变分公式精确控制。
The Shannon Noiseless coding theorem (the data-compression principle) asserts that for an information source with an alphabet $\mathcal X=\{0,\ldots ,\ell -1\}$ and an asymptotic equipartition property, one can reduce the number of stored strings $(x_0,\ldots ,x_{n-1})\in {\mathcal X}^n$ to $\ell^{nh}$ with an arbitrary small error-probability. Here $h$ is the entropy rate of the source (calculated to the base $\ell$). We consider further reduction based on the concept of utility of a string measured in terms of a rate of a weight function. The novelty of the work is that the distribution of memory is analyzed from a probabilistic point of view. A convenient tool for assessing the degree of reduction is a probabilistic large deviation principle. Assuming a Markov-type setting, we discuss some relevant formulas, including the case of a general alphabet.
研究动机与目标
- 通过在数据压缩中引入基于效用的字符串选择,扩展香农无噪声编码定理。
- 将数据存储建模为约束优化问题,仅存储由效用函数加权的“有价值”字符串。
- 使用大偏差理论分析马尔可夫源和独立同分布源的渐近存储体积。
- 基于熵和效用约束,推导存储减少的精确渐近速率。
- 为超越标准熵基方法的数据压缩中的内存使用优化提供概率框架。
提出的方法
- 使用权重函数 $\phi_n(\mathbf{x}_{0}^{n-1})$ 为字符串 $\mathbf{x}_{0}^{n-1} \in \mathcal{X}^n$ 分配效用,区分加法形式与乘法形式。
- 应用大偏差理论分析在概率测度 $p^{\rm{st}}_n$ 下所选字符串体积的渐近行为。
- 定义一个约束集合 $\mathcal{B}_n$,其中字符串的效用超过阈值 $\eta$,且信息含量不超过 $h + \epsilon$,其中 $h$ 为熵率。
- 将渐近存储速率 $\gamma(\epsilon,\eta)$ 定义为在满足效用和熵约束的测度集合上,熵泛函 $H(\upsilon)$ 的下确界。
- 对于乘法权重函数 $\phi_n(\mathbf{x}) = \prod_{i=0}^{n-1} \psi(x_i)$,在变分公式中用 $\log \psi$ 替代 $\varphi$。
- 利用大偏差原理刻画典型集概率的指数衰减速率,将其与熵率和效用约束联系起来。
实验结果
研究问题
- RQ1当字符串的选择不仅基于概率,还基于效用函数时,如何优化数据压缩?
- RQ2存储所有效用超过给定阈值 $\eta$ 的字符串所需的渐近存储体积是多少?
- RQ3大偏差原理如何允许在效用约束下推导存储减少的精确渐近速率?
- RQ4熵率 $h$ 和权重函数 $\varphi$ 在决定可实现的最小存储体积中起什么作用?
- RQ5在独立同分布或马尔可夫源的情况下,存储速率的变分公式能否简化?
主要发现
- 效用至少为 $\eta$ 且信息含量至多为 $h + \epsilon$ 的字符串的渐近存储速率由 $\gamma(\epsilon, \eta) = \inf \{ H(\upsilon) : \upsilon \in A \} $ 给出,其中 $A$ 是满足效用和熵约束的测度集合。
- 对于独立同分布源,存储速率简化为 $\gamma(\epsilon, \eta) = \inf \{ H(\upsilon) : \upsilon \in D \} $,其中 $D$ 由 $\varphi = \log \psi$ 的积分约束和 $-\int \log p^{\rm{so}}(x) \, \upsilon(dx) \leq h + \epsilon$ 定义。
- 对于所选字符串的体积 $v_n = \nu^n(\mathcal{B}_n)$,有 $\lim_{n \to \infty} \frac{1}{n} \log v_n = \gamma({\tt P}^{\rm{so}}, \epsilon, \eta)$,确认了存储减少的指数速率。
- 对于乘法权重函数 $\phi_n(\mathbf{x}) = \prod \psi(x_i)$,存储速率为 $\iota({\tt P}^{\rm{so}}, \epsilon, \eta)$,其变分公式相同,但 $\varphi$ 被替换为 $\log \psi$。
- 熵泛函 $H(\mu)$ 是凹的,约束集 $D$ 是凸的,表明在某些条件下最小化器可能唯一。
- 该框架通过引入效用,推广了香农无噪声编码定理,为选择性数据压缩提供了有理论保证的渐近高效优化方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。