Skip to main content
QUICK REVIEW

[论文解读] Database Reconstruction Is Not So Easy and Is Different from Reidentification

Krishnamurty Muralidhar, Josep Domingo‐Ferrer|arXiv (Cornell University)|Jan 24, 2023
Privacy-Preserving Technologies in Data被引用 4
一句话总结

本文挑战了广泛持有的观点,即发布统计数据必然导致数据库重建,认为传统的统计披露控制(SDC)技术——尤其是以地理层级发布时——可在不造成差分隐私中所见严重效用损失的情况下,有效防止重建。此外,本文警告称,将重建准确率作为再识别风险的代理指标,会导致对隐私风险的夸大评估。

ABSTRACT

In recent years, it has been claimed that releasing accurate statistical information on a database is likely to allow its complete reconstruction. Differential privacy has been suggested as the appropriate methodology to prevent these attacks. These claims have recently been taken very seriously by the U.S. Census Bureau and led them to adopt differential privacy for releasing U.S. Census data. This in turn has caused consternation among users of the Census data due to the lack of accuracy of the protected outputs. It has also brought legal action against the U.S. Department of Commerce. In this paper, we trace the origins of the claim that releasing information on a database automatically makes it vulnerable to being exposed by reconstruction attacks and we show that this claim is, in fact, incorrect. We also show that reconstruction can be averted by properly using traditional statistical disclosure control (SDC) techniques. We further show that the geographic level at which exact counts are released is even more relevant to protection than the actual SDC method employed. Finally, we caution against confusing reconstruction and reidentification: using the quality of reconstruction as a metric of reidentification results in exaggerated reidentification risk figures.

研究动机与目标

  • 挑战一种假设,即发布统计数据必然导致原始数据库的重建。
  • 证明传统SDC技术可有效防止重建攻击。
  • 强调地理层级在数据发布中对隐私保护的关键作用。
  • 警告不要将重建与再识别混淆,指出重建指标会夸大再识别风险。
  • 呼吁对差分隐私之外的替代隐私保护方法进行独立、同行评审的评估。

提出的方法

  • 分析数据库重建攻击的理论与实践基础,特别是Dinur-Nissim定理。
  • 比较传统SDC方法(如数据掩蔽、抑制、扰动)在防止重建方面的有效性。
  • 评估地理粒度——特别是精确计数的发布层级——对重建风险的影响。
  • 研究美国2020年人口普查中使用差分隐私(DP)的情况,包括使用拉普拉斯分布和离散高斯分布添加噪声。
  • 分析美国人口普查DAS系统中隐私参数(ε)的演变,从ε=4.5到ε=39.907,及其对隐私损失的影响。
  • 评估DAS发布的人口普查2021年和2020年数据中的实际数据不一致性,如负数家庭数量和不可能的人口分布。

实验结果

研究问题

  • RQ1声称发布统计数据必然导致数据库重建是否准确?还是基于有缺陷的假设?
  • RQ2传统统计披露控制技术是否能在不依赖差分隐私的情况下,有效防止数据库重建?
  • RQ3数据发布的地理层级在多大程度上影响重建风险?
  • RQ4将重建准确率作为再识别风险的代理指标,会在多大程度上导致对隐私威胁的高估?
  • RQ5鉴于其严重的效用和一致性问题,差分隐私是否是保护美国人口普查数据的最优方法?

主要发现

  • 声称发布统计数据可导致数据库重建的说法是错误的,其基于不完整且不透明的比较。
  • 传统SDC技术,尤其是当应用于适当的地理层级时,足以防止重建攻击。
  • 发布精确计数的地理层级,比所采用的具体SDC方法对隐私保护更为关键。
  • 将重建准确率作为再识别风险的代理指标,会导致对隐私威胁的高估,因为重建与再识别本质上是不同的攻击类型。
  • 美国人口普查局使用ε=39.907的差分隐私,与ε=4.5相比,隐私损失因子超过2.38×10^15,导致其无法提供有意义的隐私保护。
  • DAS发布的人口普查2021年和2020年数据存在严重不一致,如负数家庭数量和不可能的人口分布,表明差分隐私未能确保数据的一致性和效用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。