Skip to main content
QUICK REVIEW

[论文解读] Source Coding, Large Deviations, and Approximate Pattern Matching

Amir Dembo, Ioannis Kontoyiannis|ArXiv.org|Mar 2, 2001
Algorithms and Data Compression参考文献 79被引用 6
一句话总结

本文利用大偏差理论,提出了一种损失性版本的渐近均分性大数定律(AEP),以统一并推广率失真理论与用于有损数据压缩的模式匹配算法。该文为独立同分布及平稳过程建立了广义AEP,证明了第二阶编码定理,并刻画了等待时间、匹配长度与球覆盖指数,为分析Lempel-Ziv型有损压缩方案及失配码书提供了严格的理论基础,精确描述了其渐近行为。

ABSTRACT

We present a development of parts of rate-distortion theory and pattern- matching algorithms for lossy data compression, centered around a lossy version of the Asymptotic Equipartition Property (AEP). This treatment closely parallels the corresponding development in lossless compression, a point of view that was advanced in an important paper of Wyner and Ziv in 1989. In the lossless case we review how the AEP underlies the analysis of the Lempel-Ziv algorithm by viewing it as a random code and reducing it to the idealized Shannon code. This also provides information about the redundancy of the Lempel-Ziv algorithm and about the asymptotic behavior of several relevant quantities. In the lossy case we give various versions of the statement of the generalized AEP and we outline the general methodology of its proof via large deviations. Its relationship with Barron's generalized AEP is also discussed. The lossy AEP is applied to: (i) prove strengthened versions of Shannon's source coding theorem and universal coding theorems; (ii) characterize the performance of mismatched codebooks; (iii) analyze the performance of pattern- matching algorithms for lossy compression; (iv) determine the first order asymptotics of waiting times (with distortion) between stationary processes; (v) characterize the best achievable rate of weighted codebooks as an optimal sphere-covering exponent. We then present a refinement to the lossy AEP and use it to: (i) prove second order coding theorems; (ii) characterize which sources are easier to compress; (iii) determine the second order asymptotics of waiting times; (iv) determine the precise asymptotic behavior of longest match-lengths. Extensions to random fields are also given.

研究动机与目标

  • 将渐近均分性大数定律(AEP)扩展至在失真约束下的有损压缩,通过大偏差理论建立广义AEP。
  • 通过随机编码与典型集合论证,为分析有损模式匹配算法(包括Lempel-Ziv方案)提供理论基础。
  • 推导有损源编码、等待时间与匹配长度的第二阶渐近行为,以改进一阶结果。
  • 以最优球覆盖指数表征失配码书与加权码书的性能。
  • 将框架扩展至随机场,为空间过程建立一阶与第二阶结果。

提出的方法

  • 利用大偏差理论,为有损压缩形式化广义AEP,证明对数概率密度收敛于熵失真率。
  • 将广义AEP应用于随机码与失配码书,推导误差概率与冗余的界。
  • 利用Borel-Cantelli引理与指示变量的方差界,分析平稳过程中的等待时间与匹配长度。
  • 通过高阶渐近展开细化广义AEP,推导第二阶编码定理。
  • 引入球覆盖指数公式,表征加权失真下最优码书设计。
  • 通过空间混合条件与基于格子的采样,将结果扩展至随机场,推广等待时间与匹配长度的分析。

实验结果

研究问题

  • RQ1如何在失真约束下将渐近均分性大数定律广义化至有损压缩?
  • RQ2有损源编码速率的第二阶渐近行为是什么?它如何改进香农的直接与逆定理?
  • RQ3平稳过程间的等待时间如何随失真变化?其与模式匹配算法有何关联?
  • RQ4Lempel-Ziv等有损压缩方案中,最长匹配长度的精确渐近行为是什么?
  • RQ5失配码书如何影响压缩性能?加权码书的最优球覆盖指数是什么?

主要发现

  • 广义AEP在独立同分布与平稳过程下成立,对数概率密度以概率收敛于率失真函数。
  • 建立了第二阶有损源编码定理,表明冗余量级为 $\Theta(\sqrt{n})$(对独立同分布源)。
  • 具有失真 $D$ 的平稳过程间的等待时间满足 $\log W_n \sim \log n$ 几乎必然成立,其精确渐近界通过Borel-Cantelli引理导出。
  • 在 $d$-维随机场中,平稳过程间的最长匹配长度满足 $W_n \sim n^{1/d}$,其对数增长具有紧致界。
  • 加权码书的最优球覆盖指数被表征为关于失真度量的变分问题的下确界。
  • 对于失配码书,本文推导出可达速率的逆向界,其依赖于真实分布与失配分布之间的散度。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。