Skip to main content
QUICK REVIEW

[论文解读] A Novel Microdata Privacy Disclosure Risk Measure

Marmar Orooji, Gerald M. Knapp|arXiv (Cornell University)|Jan 2, 2019
Privacy-Preserving Technologies in Data参考文献 5被引用 12
一句话总结

本文提出了一种新颖的微观数据隐私披露风险度量方法,该方法在统一框架下联合量化身份披露风险与属性披露风险,克服了以往方法将这两种风险分开处理且假设准标识符与敏感属性之间存在严格区分的局限性。该方法采用灵活的攻击者模型和高效算法,在真实的社会工作数据集上进行了验证,展示了其可行性与性能。

ABSTRACT

A tremendous amount of individual-level data is generated each day, of use to marketing, decision makers, and machine learning applications. This data often contain private and sensitive information about individuals, which can be disclosed by adversaries. An adversary can recognize the underlying individual's identity for a data record by looking at the values of quasi-identifier attributes, known as identity disclosure, or can uncover sensitive information about an individual through attribute disclosure. In Statistical Disclosure Control, multiple disclosure risk measures have been proposed. These share two drawbacks: they do not consider identity and attribute disclosure concurrently in the risk measure, and they make restrictive assumptions on an adversary's knowledge by assuming certain attributes are quasi-identifiers and there is a clear boundary between quasi-identifiers and sensitive information. In this paper, we present a novel disclosure risk measure that addresses these limitations, by presenting a single combined metric of identity and attribute disclosure risk, and providing flexibility in modeling adversary's knowledge. We have developed an efficient algorithm for computing the proposed risk measure and evaluated the feasibility and performance of our approach on a real-world data set from the domain of social work.

研究动机与目标

  • 解决现有披露风险度量方法将身份披露与属性披露视为独立风险的局限性。
  • 消除对哪些属性为准标识符或敏感信息的刚性假设。
  • 开发一种统一且灵活的风险度量指标,以更真实地建模攻击者的知识。
  • 设计并实现一种高效算法,用于计算所提出的风险度量。
  • 在社会工作领域的真实微观数据集上评估该方法的可行性和性能。

提出的方法

  • 提出一种综合披露风险度量,将身份披露风险与属性披露风险整合为单一指标。
  • 通过允许任意属性均可对披露风险产生贡献,灵活建模攻击者知识,无需预先定义准标识符状态。
  • 采用概率框架估算重新识别和敏感信息泄露的可能性。
  • 采用计算高效的算法,使风险计算可扩展至真实世界数据集。
  • 将该风险度量应用于来自社会工作领域的实际微观数据集,以证明其实际可行性。
  • 采用基于阈值的方法评估披露风险等级,从而支持可操作的隐私决策。

实验结果

研究问题

  • RQ1如何将身份披露风险与属性披露风险统一为一个连贯的单一风险度量?
  • RQ2灵活的攻击者模型在多大程度上能提升隐私风险评估的真实性?
  • RQ3所提出的风险度量能否在真实世界微观数据集上实现高效计算?
  • RQ4与现有披露风险度量相比,该方法在准确性和灵活性方面表现如何?
  • RQ5在真实社会工作数据环境中部署该风险度量的实际可行性如何?

主要发现

  • 所提出的度量方法成功地将身份披露风险与属性披露风险统一为一个连贯的单一指标。
  • 该方法消除了对预定义准标识符边界的依赖,使攻击者知识的建模更加真实。
  • 算法在处理真实世界微观数据集方面表现出计算可行性。
  • 在社会工作数据集上的评估证实了该方法的实际适用性与性能表现。
  • 与传统方法相比,该方法能够实现更细致且准确的隐私风险评估。
  • 结果表明,该风险度量可有效用于指导隐私敏感领域中的数据发布决策。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。