[Paper Review] On-Average KL-Privacy and its equivalence to Generalization for Max-Entropy Mechanisms
This paper introduces On-Average KL-Privacy, a weakened privacy notion that is equivalent to generalization for max-entropy mechanisms—distributions proportional to exp(−ℒ(output, data))π(output). It establishes that for this broad class of algorithms, including Bayesian inference and empirical risk minimization, On-Average KL-Privacy characterizes generalization error, offering a new information-theoretic link between privacy and generalization with favorable utility trade-offs compared to differential privacy.
We define On-Average KL-Privacy and present its properties and connections to differential privacy, generalization and information-theoretic quantities including max-information and mutual information. The new definition significantly weakens differential privacy, while preserving its minimalistic design features such as composition over small group and multiple queries as well as closeness to post-processing. Moreover, we show that On-Average KL-Privacy is **equivalent** to generalization for a large class of commonly-used tools in statistics and machine learning that samples from Gibbs distributions---a class of distributions that arises naturally from the maximum entropy principle. In addition, a byproduct of our analysis yields a lower bound for generalization error in terms of mutual information which reveals an interesting interplay with known upper bounds that use the same quantity.
Motivation & Objective
- To address the gap between strong differential privacy and practical utility by proposing a weaker privacy notion that still ensures generalization.
- To formalize a connection between privacy and generalization for a broad class of statistical and machine learning algorithms that sample from Gibbs distributions.
- To demonstrate that On-Average KL-Privacy is both necessary and sufficient for generalization in max-entropy mechanisms.
- To provide a framework that weakens differential privacy while preserving key properties like composition and post-processing invariance.
- To reveal a new lower bound on generalization error in terms of mutual information, complementing known upper bounds.
Proposed method
- Proposes On-Average KL-Privacy as a relaxation of differential privacy, defined as the expected KL divergence between output distributions under neighboring datasets.
- Analyzes algorithms that sample from max-entropy distributions of the form p(h|Z) ∝ exp(−ℒ(h,Z))π(h), common in Bayesian inference and ERM.
- Uses information-theoretic tools, including mutual information I(𝒜(Z);Z) and max-information I∞(𝒜,n), to characterize privacy and generalization.
- Applies adaptive composition theorems to derive bounds on On-Average KL-Privacy under sequential queries.
- Employs Jensen’s inequality and expectation over data perturbations to relate On-Average KL-Privacy to generalization error.
- Derives a lower bound on generalization error using mutual information, contrasting with known upper bounds using the same quantity.
Experimental results
Research questions
- RQ1Can a weak privacy notion be equivalent to generalization for a broad class of learning algorithms?
- RQ2How does On-Average KL-Privacy relate to differential privacy and generalization in max-entropy mechanisms?
- RQ3What is the role of mutual information in characterizing generalization error for posterior sampling algorithms?
- RQ4Can On-Average KL-Privacy be composed adaptively across multiple queries while preserving privacy guarantees?
- RQ5What is the interplay between mutual information and generalization error in the context of max-entropy sampling?
Key findings
- On-Average KL-Privacy is equivalent to generalization for all max-entropy mechanisms, including Bayesian inference and empirical risk minimization.
- The paper establishes a lower bound on generalization error in terms of mutual information, revealing a dual role of mutual information in both upper and lower bounds.
- On-Average KL-Privacy inherits key properties of differential privacy, including composition over multiple queries and invariance under post-processing.
- The equivalence between On-Average KL-Privacy and generalization holds specifically for algorithms that sample from Gibbs distributions with a loss function and prior.
- The adaptive composition result shows that On-Average KL-Privacy is preserved under sequential queries, with a bound scaling linearly with the number of queries.
- The analysis reveals that the expected KL divergence between posterior distributions under neighboring datasets characterizes generalization error for max-entropy mechanisms.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.