Skip to main content
QUICK REVIEW

[论文解读] Structure and Sensitivity in Differential Privacy: Comparing K-Norm Mechanisms

Jordan Awan, Aleksandra Slavković|arXiv (Cornell University)|Jan 28, 2018
Privacy-Preserving Technologies in Data参考文献 46被引用 9
一句话总结

本文提出了一套框架,通过将敏感性空间的凸包识别为最优范数,优化了差分隐私中的K-范数机制,显著提升了有限样本统计推断的实用性。该研究证明,采用此方法可将所需的隐私预算减少多达一半,同时在线性回归和逻辑回归应用中保持相同的准确度。

ABSTRACT

Differential privacy (DP), provides a framework for provable privacy protection against arbitrary adversaries, while allowing the release of summary statistics and synthetic data. We address the problem of releasing a noisy real-valued statistic vector $T$, a function of sensitive data under DP, via the class of $K$-norm mechanisms with the goal of minimizing the noise added to achieve privacy. First, we introduce the sensitivity space of $T$, which extends the concepts of sensitivity polytope and sensitivity hull to the setting of arbitrary statistics $T$. We then propose a framework consisting of three methods for comparing the $K$-norm mechanisms: 1) a multivariate extension of stochastic dominance, 2) the entropy of the mechanism, and 3) the conditional variance given a direction, to identify the optimal $K$-norm mechanism. In all of these criteria, the optimal $K$-norm mechanism is generated by the convex hull of the sensitivity space. Using our methodology, we extend the objective perturbation and functional mechanisms and apply these tools to logistic and linear regression, allowing for private releases of statistical results. Via simulations and an application to a housing price dataset, we demonstrate that our proposed methodology offers a substantial improvement in utility for the same level of risk.

研究动机与目标

  • 解决有限样本设置下差分隐私统计分析中隐私与效用之间的权衡问题。
  • 开发一种系统性方法,用于比较标准L1/L2范数之外的K-范数机制。
  • 识别在保持隐私保证的同时最小化噪声的最优K-范数机制。
  • 通过敏感性空间分析,扩展功能扰动和目标扰动机制,用于私有回归。
  • 通过降低实现相同准确度所需的隐私损失预算ε,提升差分隐私的实际可用性。

提出的方法

  • 引入统计量T的敏感性空间,将敏感性多面体和凸包概念推广至任意实值统计量。
  • 提出三种比较标准:多变量随机优势、机制熵和各方向的条件方差。
  • 建立最优K-范数机制由敏感性空间的凸包生成的结论。
  • 将该框架应用于私有线性回归和逻辑回归的功能扰动与目标扰动机制。
  • 使用蒙特卡洛模拟和真实住房数据,在不同K-范数下评估性能表现。
  • 证明对统计量进行等比例缩放可缓解方向性方差失衡问题,从而避免性能排名出现反转。

实验结果

研究问题

  • RQ1对于给定统计量T,哪种K-范数机制能在满足差分隐私的前提下最小化噪声?
  • RQ2如何利用多变量随机优势和熵来比较K-范数机制?
  • RQ3敏感性空间的凸包是否在所有损失函数下均一致产生最优K-范数机制?
  • RQ4所提出的框架能否在保持统计准确度的前提下,减少回归任务中所需的隐私损失预算ε?
  • RQ5在何种条件下,更大的范数球H可能在边际方差或损失方面优于更小的范数球K,即使其条件方差更高?

主要发现

  • 敏感性空间的凸包在所有三种比较标准下(随机优势、熵和条件方差)均一致生成最优K-范数机制。
  • 在线性回归和逻辑回归中,所提出的方法可在约一半的标准机制隐私损失预算ε下实现相同准确度。
  • 在反例中,尽管K范数在每个方向上的条件方差更低,但更大的范数球H在期望ℓ₁、ℓ₂和ℓ∞损失方面仍优于K范数机制。
  • 这种性能反转源于变量间缩放的不均衡,揭示了效用度量中的多变量“辛普森悖论”。
  • 当变量被等比例缩放时,基于凸包的K-范数机制在所有效用度量中均持续优于其他方法。
  • 该框架显著提升了有限样本差分隐私统计推断的实际效用,尤其在回归和合成数据生成任务中。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。