Skip to main content
QUICK REVIEW

[Paper Review] Strong Data Processing Inequalities for Input Constrained Additive Noise Channels

Flávio P. Calmon, Yury Polyanskiy|arXiv (Cornell University)|Dec 20, 2015
Wireless Communication Security Techniques19 references3 citations
TL;DR

This paper establishes strong data processing inequalities (SDPIs) for additive noise channels with input constraints, proving that mutual information $I(W;Y)$ is strictly less than $I(W;X)$ for any non-degenerate noise with a density whose support is not disjoint from its translates. The key result shows that for such channels, $I(W;Y) \leq F_I(I(W;X))$ with $F_I(t) < t$ for all $t > 0$, and that capacity is only approached when $I(W;X) \to \infty$, under quadratic power constraints. Explicit, order-optimal bounds are derived for the Gaussian channel using I-MMSE and Talagrand’s inequality.

ABSTRACT

This paper quantifies the intuitive observation that adding noise reduces available information by means of non-linear strong data processing inequalities. Consider the random variables $W o X o Y$ forming a Markov chain, where $Y=X+Z$ with $X$ and $Z$ real-valued, independent and $X$ bounded in $L_p$-norm. It is shown that $I(W;Y) \le F_I(I(W;X))$ with $F_I(t)0$, if and only if $Z$ has a density whose support is not disjoint from any translate of itself. A related question is to characterize for what couplings $(W,X)$ the mutual information $I(W;Y)$ is close to maximum possible. To that end we show that in order to saturate the channel, i.e. for $I(W;Y)$ to approach capacity, it is mandatory that $I(W;X) o\infty$ (under suitable conditions on the channel). A key ingredient for this result is a deconvolution lemma which shows that post-convolution total variation distance bounds the pre-convolution Kolmogorov-Smirnov distance. Explicit bounds are provided for the special case of the additive Gaussian noise channel with quadratic cost constraint. These bounds are shown to be order-optimal. For this case simplified proofs are provided leveraging Gaussian-specific tools such as the connection between information and estimation (I-MMSE) and Talagrand's information-transportation inequality.

Motivation & Objective

  • To quantify how much mutual information is lost when passing through an additive noise channel with input constraints.
  • To establish non-linear strong data processing inequalities (SDPIs) that strictly contract mutual information for non-degenerate noise.
  • To characterize when mutual information $I(W;Y)$ approaches the channel capacity, showing it requires $I(W;X) \to \infty$ under suitable conditions.
  • To derive explicit, order-optimal bounds for the additive Gaussian noise channel with quadratic power constraints.

Proposed method

  • The authors define a data-processing function $F_I(t)$ that maps input mutual information $I(W;X)$ to the maximum possible output mutual information $I(W;Y)$, under input constraints.
  • They prove that $F_I(t) < t$ for all $t > 0$ if and only if the noise $Z$ has a density whose support is not disjoint from any of its translates.
  • A key technical tool is a deconvolution lemma that bounds the pre-convolution Kolmogorov-Smirnov distance using post-convolution total variation distance.
  • For Gaussian channels, the authors leverage the I-MMSE relationship and Talagrand’s information-transportation inequality to derive simplified, order-optimal bounds.
  • They analyze the infinite-dimensional case and prove that $F_I^\infty(t,\gamma)/t$ is decreasing and $F_I^\infty(t,\gamma)$ is increasing, enabling tight asymptotic bounds.
  • The proof of the main result uses characteristic function bounds and Esseen’s inequality to control the distance between the input distribution and a Gaussian.

Experimental results

Research questions

  • RQ1Under what conditions on the noise $Z$ does mutual information strictly contract under an additive noise channel with input constraints?
  • RQ2What is the precise functional relationship $F_I(t)$ such that $I(W;Y) \leq F_I(I(W;X))$ for all $W \to X \to Y$ with $Y = X + Z$?
  • RQ3Can the mutual information $I(W;Y)$ approach the channel capacity only if $I(W;X) \to \infty$ under quadratic power constraints?
  • RQ4How can explicit, order-optimal bounds on $F_I(t)$ be derived for the additive Gaussian noise channel with input power constraints?
  • RQ5What role does the Lévy concentration function play in characterizing the absence of atoms in the input distribution?

Key findings

  • The mutual information $I(W;Y)$ is strictly less than $I(W;X)$ for all $t > 0$ if and only if the noise $Z$ has a density whose support is not disjoint from any of its translates.
  • For the additive Gaussian noise channel with quadratic power constraint $\gamma$, the function $F_I^\infty(t,\gamma)$ satisfies $F_I^\infty(\gamma/2,\gamma) = \gamma/2$, showing that the optimal rate is achieved only in the limit of infinite input mutual information.
  • The bound $F_I(t) < t$ holds uniformly for all $t > 0$ under the stated condition on the noise, proving strict data processing contraction.
  • Using Talagrand’s inequality and I-MMSE tools, the paper derives explicit bounds on the Kolmogorov-Smirnov distance between the input and a Gaussian, showing near-Gaussianity under high mutual information.
  • The deconvolution lemma establishes that post-convolution total variation distance bounds the pre-convolution Kolmogorov-Smirnov distance, enabling the key contraction result.
  • The Lévy concentration function $\mathcal{L}(X;\delta)$ is continuous at zero if and only if the distribution has no atoms, which is used to analyze the behavior of the input distribution near the Gaussian limit.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.