Skip to main content
QUICK REVIEW

[Paper Review] Information Theory of Data Privacy

Genqiang Wu, Xianyao Xia|arXiv (Cornell University)|Mar 22, 2017
Privacy-Preserving Technologies in Data35 references3 citations
TL;DR

This paper proposes a Bayesian inference-based privacy model that integrates Shannon's cryptography framework with a lower bound on adversary uncertainty to quantify privacy loss. By ensuring adversaries gain minimal information per individual when their uncertainty exceeds this bound—especially in large datasets—it provides a principled justification for using Bayesian risk factors in privacy metrics, while enabling flexible trade-offs between privacy and utility through four adjustable parameters.

ABSTRACT

By combining Shannon's cryptography model with an assumption to the lower bound of adversaries' uncertainty to the queried dataset, we develop a secure Bayesian inference-based privacy model and then in some extent answer Dwork et al.'s question [1]: why Bayesian risk factors are the right measure for privacy loss. This model ensures an adversary can only obtain little information of each individual from the model's output if the adversary's uncertainty to the queried dataset is larger than the lower bound. Importantly, the assumption to the lower bound almost always holds, especially for big datasets. Furthermore, this model is flexible enough to balance privacy and utility: by using four parameters to characterize the assumption, there are many approaches to balance privacy and utility and to discuss the group privacy and the composition privacy properties of this model.

Motivation & Objective

  • To address Dwork et al.'s open question on why Bayesian risk factors are appropriate for measuring privacy loss.
  • To formalize a privacy model grounded in information theory that limits adversary knowledge about individuals in a dataset.
  • To establish a lower bound on adversary uncertainty that is practically always satisfied, especially in large datasets.
  • To enable tunable balance between privacy and utility through four configurable parameters.
  • To support analysis of group privacy and composition properties within the model framework.

Proposed method

  • Integrates Shannon's cryptographic model with Bayesian inference to model privacy as a function of adversary uncertainty.
  • Introduces a lower bound on the adversary's uncertainty about the queried dataset as a foundational assumption.
  • Uses four parameters to characterize the uncertainty assumption, enabling flexible control over privacy-utility trade-offs.
  • Derives privacy loss as a function of Bayesian risk, showing that low information gain occurs when adversary uncertainty exceeds the lower bound.
  • Applies the model to analyze group privacy and composition properties, demonstrating its scalability and robustness.
  • Ensures theoretical guarantees by relying on the assumption that the lower bound on uncertainty holds—particularly valid for big datasets.

Experimental results

Research questions

  • RQ1Why are Bayesian risk factors the appropriate measure for quantifying privacy loss in data publishing?
  • RQ2How can a privacy model be constructed that ensures minimal individual information leakage when adversary uncertainty is bounded below?
  • RQ3What are the implications of this model for group privacy and composition of multiple queries?
  • RQ4In what way can the model flexibly balance privacy and utility through parameterization?
  • RQ5Under what conditions does the lower bound on adversary uncertainty hold, and how does this affect the model’s validity?

Key findings

  • The model justifies the use of Bayesian risk factors as a natural measure for privacy loss by linking it to the adversary’s uncertainty about the dataset.
  • When adversary uncertainty exceeds the lower bound, the model ensures that each individual’s information is minimally exposed in the output.
  • The lower bound assumption holds almost always in large datasets, making the model practically viable and robust.
  • The model supports flexible privacy-utility trade-offs through four adjustable parameters, enabling customization for different threat models.
  • The framework naturally accommodates analysis of group privacy and composition of queries, enhancing its applicability to real-world systems.
  • The theoretical foundation ensures that privacy loss is bounded and predictable under realistic assumptions about adversary knowledge.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.