Skip to main content
QUICK REVIEW

[论文解读] Bias and Variance of Post-processing in Differential Privacy

Kèyù Zhü, Pascal Van Hentenryck|arXiv (Cornell University)|Oct 9, 2020
Privacy-Preserving Technologies in Data参考文献 11被引用 6
一句话总结

本文研究了差分隐私中后处理的统计影响,重点关注用于强制执行人口普查数据领域约束的投影。它从理论和实证两方面分析了后处理差分隐私输出中的偏差与方差,表明非负性约束会引入偏差,其偏差受推导公式的有界限制;同时,随着维度增加,残差误差收敛于拉普拉斯分布,通过利用公共信息降低方差。

ABSTRACT

Post-processing immunity is a fundamental property of differential privacy: it enables the application of arbitrary data-independent transformations to the results of differentially private outputs without affecting their privacy guarantees. When query outputs must satisfy domain constraints, post-processing can be used to project the privacy-preserving outputs onto the feasible region. Moreover, when the feasible region is convex, a widely adopted class of post-processing steps is also guaranteed to improve accuracy. Post-processing has been applied successfully in many applications including census data-release, energy systems, and mobility. However, its effects on the noise distribution is poorly understood: It is often argued that post-processing may introduce bias and increase variance. This paper takes a first step towards understanding the properties of post-processing. It considers the release of census data and examines, both theoretically and empirically, the behavior of a widely adopted class of post-processing functions.

研究动机与目标

  • 理解后处理对差分隐私数据发布中噪声分布的统计影响。
  • 分析投影(常用于强制执行领域约束)对隐私保护输出中偏差与方差的影响。
  • 为层次化数据发布任务中后处理的行为提供理论与实证见解,特别是在人口普查应用中的表现。
  • 表征后处理后的残差分布,并评估其对准确性与公平性的影响。

提出的方法

  • 本文研究了两类投影:一类无非负性约束,另一类有非负性约束,均在满足线性可行性约束的条件下进行。
  • 利用顺序统计量与拉普拉斯噪声的性质,推导出在存在非负性约束时,投影引入偏差的上界。
  • 将残差误差建模为独立同分布的拉普拉斯随机变量的线性组合,并分析其极限分布。
  • 应用斯利茨基定理与大数定律,证明当维度 n → ∞ 时,残差分布收敛于拉普拉斯分布。
  • 在真实人口普查数据(亚利桑那州与德克萨斯州)上进行实证验证,通过不同县数量评估方差与偏差。
  • 采用一个包含单一线性方程的子问题,完全表征了在层次化数据约束下后处理噪声分布的特性。

实验结果

研究问题

  • RQ1使用投影进行后处理是否会在差分隐私输出中引入偏差?若会,其条件是什么?
  • RQ2非负性约束的存在如何影响后处理引入的偏差?
  • RQ3后处理后残差误差的极限分布是什么?其与原始噪声分布有何关系?
  • RQ4后处理如何影响输出的方差,特别是在高维设置下?
  • RQ5后处理在多大程度上提升了层次化数据发布任务中的准确性?

主要发现

  • 当不存在非负性约束时,通过投影进行的后处理不会引入偏差。
  • 在存在非负性约束时,本文推导出偏差的上界,有助于识别偏差将显著的问题类型。
  • 后处理输出的残差误差在维度 n 增加时,其分布收敛于原始拉普拉斯噪声。
  • 后处理数据的方差随维度增加而减小,表明通过利用公共信息可提升准确性。
  • 在亚利桑那州与德克萨斯州人口普查数据上的实证结果表明,理论方差与实证方差高度一致(例如,亚利桑那州为 186.67 与 186.88)。
  • 残差分布收敛于拉普拉斯分布,表明后处理可在保持隐私的同时降低方差,对群体公平性具有重要意义。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。