Skip to main content
QUICK REVIEW

[论文解读] On the Tradeoff between Privacy and Distortion in Differential Privacy.

Weina Wang, Lei Ying|arXiv (Cornell University)|Feb 16, 2014
Privacy-Preserving Technologies in Data参考文献 5被引用 5
一句话总结

本文通过引入隐私-失真函数 ∗(D),建立了在差分隐私合成数据库生成中的基本隐私-失真权衡。该函数量化了在失真约束下可实现的最小差分隐私级别。本文提出了一种计算高效的机制 E,在均匀先验下实现最优性,并揭示了在一种新型后验差分隐私概念下,隐私-失真与信息论中的率失真理论之间的深刻联系。

ABSTRACT

In this paper, we consider the setting in which the output of a differentially private mechanism is in the same universe as the input, and investigate the usefulness in terms of (the negative of) the distortion between the output and the input. This setting can be regarded as the synthetic database release problem. We define a privacy–distortion function ∗(D), which is the smallest (best) achievable differential privacy level given a distortion upper bound D, and quantify the fundamental privacy–distortion tradeoff by characterizing ∗. Specifically, we first obtain an upper bound on ∗ by designing a mechanism E. Then we derive a lower bound on ∗ that deviates from the upper bound only by a constant. It turns out that E is an optimal mechanism when the database is drawn uniformly from the universe, i.e., the upper bound and the lower bound meet. A significant advantage of mechanism E is that its distortion guarantee does not depend on the prior and its implementation is computationally efficient, although it may not be optimal always. From a learning perspective, we further introduce a new notion of differential privacy that is defined on the posterior probabilities, which we call a posteriori differential privacy. Under this notion, the exact form of the privacy–distortion function is obtained for a wide range of distortion values. We then establish a fundamental connection between the privacy–distortion tradeoff and the information-theoretic rate–distortion theory. An interesting finding is that there exists a consistency between the rate– distortion and the privacy–distortion under a posteriori differential privacy, which is shown by devising a mechanism that minimizes the mutual information and the privacy level simultaneously. 1

研究动机与目标

  • 刻画在合成数据库发布中差分隐私与失真之间的基本权衡。
  • 定义并分析隐私-失真函数 ∗(D),该函数捕捉在给定失真约束下可实现的最小隐私级别。
  • 设计一种机制 E,实现在低计算成本下接近最优的隐私-失真权衡,并具备与先验无关的失真保证。
  • 建立隐私-失真与信息论中率失真理论之间的理论联系。
  • 引入并分析一种新的后验差分隐私概念,以实现对隐私-失真函数的精确表征。

提出的方法

  • 隐私-失真函数 ∗(D) 被定义为在失真约束 D 下可实现的最小差分隐私级别。
  • 构建一种机制 E,以提供 ∗(D) 的上界,其失真与数据先验无关。
  • 推导出 ∗(D) 的下界,表明上下界之间的差距至多为常数。
  • 证明当数据库在全集上均匀分布时,机制 E 是最优的。
  • 引入一种新的隐私概念——后验差分隐私,其定义基于后验概率而非数据分布。
  • 形式化了隐私-失真与率失真理论之间的联系,展示了一种机制可同时最小化互信息与隐私级别。

实验结果

研究问题

  • RQ1在合成数据库生成中,差分隐私与失真之间的基本权衡是什么?
  • RQ2能否设计一种差分隐私机制,其失真保证与数据先验无关?
  • RQ3隐私-失真权衡如何与经典率失真理论相关联?
  • RQ4在何种条件下,所提出的机制 E 是最优的?
  • RQ5一种新的隐私概念——后验差分隐私,能否实现对隐私-失真函数的精确表征?

主要发现

  • 机制 E 实现了接近最优的隐私-失真权衡,∗(D) 的上下界仅相差一个常数因子。
  • 当数据库在全集上均匀分布时,机制 E 是最优的,证实了其在此设定下的理论最优性。
  • 机制 E 的失真保证不依赖于数据先验,表现出对不同数据分布的鲁棒性。
  • 本文在新的后验差分隐私框架下,建立了率失真理论与隐私-失真理论之间的一致性。
  • 构建了一种机制,可同时最小化互信息与差分隐私级别,展示了信息论度量与隐私度量之间的紧密联系。
  • 在后验差分隐私框架下,对一系列失真值,隐私-失真函数 ∗(D) 实现了精确表征。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。