Skip to main content
QUICK REVIEW

[Paper Review] Mismatched Guesswork

Salman Salamatian, Litian Liu|arXiv (Cornell University)|Jul 1, 2019
Algorithms and Data Compression23 references4 citations
TL;DR

This paper introduces a large deviation principle (LDP) for mismatched guesswork, where the guesswork is evaluated using a mismatched distribution $ $ instead of the true distribution $ $. It shows that the asymptotic guesswork growth is governed by the entropy of the projection of the true distribution onto the tilted family of $ $, and that one-to-one source coding is more robust to mismatch than prefix-free coding, with performance loss vanishing if and only if the true distribution lies on the tilted family of the mismatched distribution.

ABSTRACT

We study the problem of mismatched guesswork, where we evaluate the number of symbols $y \in \mathcal{Y}$ which have higher likelihood than $X \sim μ$ according to a mismatched distribution $ν$. We discuss the role of the tilted/exponential families of the source distribution $μ$ and of the mismatched distribution $ν$. We show that the value of guesswork can be characterized using the tilted family of the mismatched distribution $ν$, while the probability of guessing is characterized by an exponential family which passes through $μ$. Using this characterization, we demonstrate that the mismatched guesswork follows a large deviation principle (LDP), where the rate function is described implicitly using information theoretic quantities. We apply these results to one-to-one source coding (without prefix free constraint) to obtain the cost of mismatch in terms of average codeword length. We show that the cost of mismatch in one-to-one codes is no larger than that of the prefix-free codes, i.e., $D(μ\| ν)$. Further, the cost of mismatch vanishes if and only if $ν$ lies on the tilted family of the true distribution $μ$, which is in stark contrast to the prefix-free codes. These results imply that one-to-one codes are inherently more robust to mismatch.

Motivation & Objective

  • To analyze the large deviations behavior of guesswork when the distribution used for guessing differs from the true source distribution.
  • To characterize the asymptotic growth rate of mismatched guesswork using geometric properties of tilted families.
  • To apply the results to one-to-one source coding and compare its robustness to mismatch with prefix-free coding.
  • To determine the conditions under which the performance cost of mismatch vanishes in source coding.

Proposed method

  • The analysis uses tilted (exponential) families of the true distribution $ $ and the mismatched distribution $ $, leveraging geometric projections onto these families.
  • The LDP rate function is implicitly characterized using relative entropy between distributions on the tilted family of $ $ and the true distribution $ $.
  • The asymptotic guesswork growth rate is derived via a variational optimization over the tilted family of $ $, linking it to the entropy of the projected distribution.
  • The connection between guesswork and one-to-one source coding is established via a correspondence between codeword length and guesswork value, enabling direct application of LDP results.
  • The reliability function and average codeword length in mismatched one-to-one coding are derived from the LDP of guesswork using L’Hôpital’s rule and limit analysis.
  • The proof relies on the fact that optimal one-to-one codes satisfy $g_{ }(x^n) \leq l(f^*_{ }(x^n)) < g_{ }(x^n)+1$, enabling translation of guesswork results to coding performance.

Experimental results

Research questions

  • RQ1How does the large deviation behavior of guesswork change when the guessing distribution is mismatched relative to the true source distribution?
  • RQ2What is the asymptotic growth rate of mismatched guesswork, and how is it characterized in terms of information-theoretic quantities?
  • RQ3How does the performance cost of mismatch in one-to-one source coding compare to that in prefix-free source coding?
  • RQ4Under what conditions does the cost of mismatch vanish in one-to-one source coding?
  • RQ5Can the reliability function and average codeword length in mismatched one-to-one coding be expressed in terms of the LDP of guesswork?

Key findings

  • The mismatched guesswork satisfies a large deviation principle (LDP) with a rate function implicitly defined via relative entropy between distributions on the tilted family of the mismatched distribution $ $ and the true distribution $ $.
  • The asymptotic growth rate of mismatched guesswork is given by $H(\Pi_{\mathcal{T}_{\nu}}(\mu))$, the entropy of the projection of $\mu$ onto the tilted family of $\n$.
  • The average codeword length in mismatched one-to-one source coding is $L(\n\|\mu) = H(\Pi_{\mathcal{T}_{\nu}}(\mu))$, which is always less than or equal to the corresponding length in prefix-free coding.
  • The cost of mismatch in one-to-one coding is bounded above by $D(\mu\|\n)$, the KL divergence between the true and mismatched distributions, and is strictly smaller than in prefix-free coding.
  • The performance loss due to mismatch vanishes if and only if $\mu$ lies on the tilted family of $\n$, i.e., $\mu \in \mathcal{T}_{\nu}^{+}$.
  • The reliability function for mismatched one-to-one coding is $E(R,\n\|\mu) = J(R)$, which matches the LDP rate function derived for guesswork, confirming the tight connection between guesswork and coding performance.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.