Skip to main content
QUICK REVIEW

[论文解读] Take it or Leave it: Running a Survey when Privacy Comes at a Cost

Katrina Ligett, Aaron Roth|arXiv (Cornell University)|Feb 21, 2012
Privacy-Preserving Technologies in Data参考文献 7被引用 12
一句话总结

本文提出了一种‘要就拿走,不要就放弃’的机制,通过向个体提供考虑隐私成本的报价来估算敏感的人口统计数据,从而绕过直接披露机制中的不可能性结果。通过将隐私成本建模为微差隐私的函数,并使用基于阈值的接受机制进行随机抽样,该方法即使在隐私成本与私人数据相关时,也能确保在有界期望成本下实现准确估计。

ABSTRACT

In this paper, we consider the problem of estimating a potentially sensitive (individually stigmatizing) statistic on a population. In our model, individuals are concerned about their privacy, and experience some cost as a function of their privacy loss. Nevertheless, they would be willing to participate in the survey if they were compensated for their privacy cost. These cost functions are not publicly known, however, nor do we make Bayesian assumptions about their form or distribution. Individuals are rational and will misreport their costs for privacy if doing so is in their best interest. Ghosh and Roth recently showed in this setting, when costs for privacy loss may be correlated with private types, if individuals value differential privacy, no individually rational direct revelation mechanism can compute any non-trivial estimate of the population statistic. In this paper, we circumvent this impossibility result by proposing a modified notion of how individuals experience cost as a function of their privacy loss, and by giving a mechanism which does not operate by direct revelation. Instead, our mechanism has the ability to randomly approach individuals from a population and offer them a take-it-or-leave-it offer. This is intended to model the abilities of a surveyor who may stand on a street corner and approach passers-by.

研究动机与目标

  • 解决当个体隐私成本与其私人数据相关时,估算敏感人口统计数据的挑战。
  • 克服直接披露机制中隐私成本与私人类型相关导致无法实现真实报告的不可能性结果。
  • 设计一种机制,确保个体理性且实现准确估计,而无需要求参与者真实报告其隐私成本。
  • 以一种可实现的方式建模隐私成本,使参与具有激励相容性,同时保持微差隐私保证。
  • 为在个体行为现实假设下兼具隐私保护与成本效率的调查机制提供一个框架。

提出的方法

  • 该机制使用随机抽样过程按顺序接触个体,为其提供‘要就拿走,不要就放弃’的补偿。
  • 每项报价基于隐私预算和微差隐私参数,接受与否由个体的私人成本所决定的阈值确定。
  • 该算法按周期运行,逐步提高报价水平并降低接受阈值,以在成本与准确性之间取得平衡。
  • 该机制采用拉普拉斯噪声机制以确保微差隐私,隐私参数在各周期间调整以控制敏感性。
  • 基于已接受集合的大小和预期参与人数,使用停止规则,确保估计值收敛至稳定状态。
  • 在各周期中分析成本,利用几何级数和概率集中不等式推导出有界结果,表明期望成本呈次线性增长。

实验结果

研究问题

  • RQ1我们能否设计一种机制,在隐私成本与私人数据相关时,实现对敏感统计数据的准确估计?
  • RQ2在直接披露机制因隐私成本相关性而失效的设定下,是否可能实现个体理性与真实报告?
  • RQ3如何建模隐私成本,以实现激励相容的参与,而无需参与者真实报告其成本?
  • RQ4此类机制的期望成本是多少?其能否相对于基准成本有界?
  • RQ5与传统直接披露方法相比,具有隐私意识报价的‘要就拿走,不要就放弃’机制在成本与准确性方面是否表现更优?

主要发现

  • 该机制实现了对总体统计数据的准确估计,其期望成本为 O(log log(α·v(α/8)) · BenchmarkCost + 1/α²),相对于基准成本呈次线性增长。
  • 当参数 η 满足 c₁ < η < 3/17 − c₂ 时,期望成本有界,其中 c₁ 和 c₂ 为远离零的常数。
  • 在 j* 之前的周期对总成本的贡献被限制为 (1+η)^j* / η · EpochSize(j*),在所选参数下该值较小。
  • 在 j* 之后的周期中,第 t 个周期仍未停止的概率至多为 (17/20)^(t−j*),确保快速收敛。
  • 该机制提供了‘近乎真实’的保证:即使参与是自愿的且严格激励,说谎的期望收益也可忽略不计。
  • 该方法通过放弃直接披露,转而采用随机化、顺序报价机制,绕过了 Ghosh 和 Roth 提出的不可能性结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。