Skip to main content
QUICK REVIEW

[论文解读] Attacks on Deidentification's Defenses

Aloni Cohen|arXiv (Cornell University)|Feb 27, 2022
Privacy-Preserving Technologies in Data被引用 5
一句话总结

本文提出了新颖的理论与实践攻击,从根本上动摇了基于准标识符的去标识化技术(如 k-匿名性)的安全性,即使所有属性均被视为准标识符。研究证明,在现实条件下,k-匿名性无法提供有意义的隐私保护,如通过对符合严格监管标准的 EdX 数据集进行去标识化后,成功重新识别出学生身份所示。

ABSTRACT

Quasi-identifier-based deidentification techniques (QI-deidentification) are widely used in practice, including $k$-anonymity, $\ell$-diversity, and $t$-closeness. We present three new attacks on QI-deidentification: two theoretical attacks and one practical attack on a real dataset. In contrast to prior work, our theoretical attacks work even if every attribute is a quasi-identifier. Hence, they apply to $k$-anonymity, $\ell$-diversity, $t$-closeness, and most other QI-deidentification techniques. First, we introduce a new class of privacy attacks called downcoding attacks, and prove that every QI-deidentification scheme is vulnerable to downcoding attacks if it is minimal and hierarchical. Second, we convert the downcoding attacks into powerful predicate singling-out (PSO) attacks, which were recently proposed as a way to demonstrate that a privacy mechanism fails to legally anonymize under Europe's General Data Protection Regulation. Third, we use LinkedIn.com to reidentify 3 students in a $k$-anonymized dataset published by EdX (and show thousands are potentially vulnerable), undermining EdX's claimed compliance with the Family Educational Rights and Privacy Act. The significance of this work is both scientific and political. Our theoretical attacks demonstrate that QI-deidentification may offer no protection even if every attribute is treated as a quasi-identifier. Our practical attack demonstrates that even deidentification experts acting in accordance with strict privacy regulations fail to prevent real-world reidentification. Together, they rebut a foundational tenet of QI-deidentification and challenge the actual arguments made to justify the continued use of $k$-anonymity and other QI-deidentification techniques.

研究动机与目标

  • 挑战一个基础假设:即当所有属性均被视为准标识符时,k-匿名性可提供有意义的隐私保护。
  • 证明即使经过专家验证并符合严格隐私法规的去标识化,仍可能在实践中被攻破。
  • 表明基于准标识符的去标识化未能满足核心隐私属性:对后处理的鲁棒性、组合下的平滑退化性,以及避免分布假设。
  • 提供理论与实证证据,表明 k-匿名性及其改进方法不足以提供法律与实践上的隐私保障。

提出的方法

  • 提出下编码攻击(downcoding attacks),一类新型隐私攻击,利用准标识符去标识化中最小化与分层泛化方案的漏洞。
  • 将下编码攻击转化为谓词单点攻击(predicate singling-out, PSO)攻击,以证明其在 GDPR 法律匿名化标准下的失败。
  • 利用 LinkedIn 数据执行真实世界的重新识别攻击,成功从一个经 k-匿名化处理的 EdX 数据集中重新识别出个体。
  • 分析 EdX 数据集,表明后处理(如行删除)可破坏 k-匿名性,违反对后处理的鲁棒性要求。
  • 通过理论分析表明,基于准标识符的去标识化中的隐私保护依赖于未明示的数据分布假设。
  • 结合理论建模与真实数据集的实证验证,从多个维度展示其脆弱性。

实验结果

研究问题

  • RQ1当所有属性均被视为准标识符时,基于准标识符的去标识化技术(如 k-匿名性)能否提供有意义的隐私保护?
  • RQ2是否存在理论与实践上的攻击,可对经专家验证并符合严格隐私法规的去标识化数据集实现个体重新识别?
  • RQ3基于准标识符的去标识化是否满足核心隐私属性,如对后处理的鲁棒性与组合下的平滑退化性?
  • RQ4是否可在最小化与分层泛化方案中构建并利用下编码攻击?
  • RQ5现实世界中的重新识别攻击在多大程度上破坏了 k-匿名性及其相关技术在 GDPR 和 FERPA 法律与监管合规性上的主张?

主要发现

  • 本文首次提出针对基于准标识符的去标识化的理论攻击——下编码攻击,即使在所有属性均被视为准标识符的情况下依然有效。
  • 一项实际的重新识别攻击成功从一个声称符合 FERPA 标准并由专家去标识化的 EdX 数据集中重新识别出三名学生。
  • 尽管 EdX 数据集经过严格监管流程处理,但其在课程属性上并非真正意义上的 5-匿名,违反了 k-匿名性的定义。
  • 这些攻击表明,基于准标识符的去标识化对后处理不具鲁棒性,因为删除行可破坏 k-匿名性,表明其定义存在语法缺陷。
  • 研究表明,基于准标识符的去标识化在组合下无法实现平滑退化,如通过多个 k-匿名化发布版本可恢复原始数据。
  • 研究结果挑战了 k-匿名性及其相关技术在法律与实践上的正当性,表明其不满足核心隐私属性,如分布鲁棒性与后处理抗性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。