Skip to main content
QUICK REVIEW

[论文解读] A Central Limit Theorem for Differentially Private Query Answering

Jinshuo Dong, Weijie Su|arXiv (Cornell University)|Mar 15, 2021
Privacy-Preserving Technologies in Data参考文献 29被引用 5
一句话总结

该论文为差分隐私查询回答建立了中心极限定理,表明在高维下,密度与 $\mathrm{e}^{-\|x\|_p^\alpha}}$ 成比例的噪声分布会收敛到高斯差分隐私(GDP)。它证明了高斯机制在隐私-精度权衡中达到最优,且隐私参数与 $\ell_2$-损失的乘积下界为维度,确认其在该类噪声家族中具有紧致最优性。

ABSTRACT

Perhaps the single most important use case for differential privacy is to privately answer numerical queries, which is usually achieved by adding noise to the answer vector. The central question, therefore, is to understand which noise distribution optimizes the privacy-accuracy trade-off, especially when the dimension of the answer vector is high. Accordingly, extensive literature has been dedicated to the question and the upper and lower bounds have been matched up to constant factors [BUV18, SU17]. In this paper, we take a novel approach to address this important optimality question. We first demonstrate an intriguing central limit theorem phenomenon in the high-dimensional regime. More precisely, we prove that a mechanism is approximately Gaussian Differentially Private [DRS21] if the added noise satisfies certain conditions. In particular, densities proportional to $\mathrm{e}^{-\|x\|_p^α}$, where $\|x\|_p$ is the standard $\ell_p$-norm, satisfies the conditions. Taking this perspective, we make use of the Cramer--Rao inequality and show an "uncertainty principle"-style result: the product of the privacy parameter and the $\ell_2$-loss of the mechanism is lower bounded by the dimension. Furthermore, the Gaussian mechanism achieves the constant-sharp optimal privacy-accuracy trade-off among all such noises. Our findings are corroborated by numerical experiments.

研究动机与目标

  • 解决高维下差分隐私数值查询回答中最佳噪声分布的根本性问题。
  • 理解为何尽管放宽了隐私约束,高斯机制在一维下仍表现不如截断拉普拉斯等替代方案。
  • 探究一维机制中的见解是否可推广至高维查询工作负载。
  • 为一般噪声机制向高斯差分隐私(GDP)的渐近收敛建立理论基础。
  • 利用Cramér–Rao不等式推导隐私-精度权衡的紧下界,表明高斯机制在常数因子内达到最优。

提出的方法

  • 证明了高维噪声添加机制的中心极限定理,表明在噪声密度的温和条件下,机制会收敛到高斯差分隐私(GDP)。
  • 分析了密度与 $\mathrm{e}^{-\|x\|_p^\alpha}}$ 成比例的噪声分布,包括拉普拉斯($p=1, \alpha=1$)和高斯($p=2, \alpha=2$)作为特例。
  • 应用Cramér–Rao不等式推导隐私参数 $\varepsilon$ 与 $\ell_2$-损失乘积的下界,表明其至少与维度 $n$ 成比例。
  • 通过对对数似然比的二阶泰勒展开分析隐私损失分布,并在高维渐近下建立向GDP收敛的结论。
  • 采用多变量版本的Lindeberg–Lévy中心极限定理,表明在适当缩放下,隐私损失随机变量收敛到正态分布。
  • 通过数值实验验证理论结果,表明在 $n=30$ 的有限维设置下,GDP收敛速度很快。

实验结果

研究问题

  • RQ1在高维渐近下,能否为差分隐私机制建立中心极限定理?
  • RQ2具有重尾或次高斯尾部的噪声分布是否会随着维度增加而收敛到高斯差分隐私?
  • RQ3在所有密度与 $\mathrm{e}^{-\|x\|_p^\alpha}}$ 成比例的噪声分布中,高斯机制是否在隐私-精度权衡上达到最优?
  • RQ4能否利用信息论工具(如Cramér–Rao不等式)推导出隐私-精度权衡的根本下界?
  • RQ5为何高斯机制在一维下表现不如截断拉普拉斯,且这种差距在高维下是否依然存在?

主要发现

  • 差分隐私机制存在中心极限定理:在噪声密度的温和正则性条件下,随着维度增加,隐私损失收敛到正态分布。
  • 密度与 $\mathrm{e}^{-\|x\|_p^\alpha}}$ 成比例的噪声分布满足向高斯差分隐私(GDP)收敛的条件,包括拉普拉斯和高斯作为特例。
  • 隐私参数 $\varepsilon$ 与机制 $\ell_2$-损失的乘积有下界,其下界为维度 $n$,从而确立了根本性的权衡。
  • 高斯机制在该下界中达到紧致常数,使其在所有此类噪声族中达到最优(至多常数因子)。
  • 数值实验表明,即使在 $n=30$ 时,GDP的收敛速度也很快,支持理论渐近结果。
  • 分析表明,尽管在一维下表现次优,但在高维下隐私-精度权衡无法超越高斯机制的性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。