[论文解读] Optimal Schemes for Discrete Distribution Estimation under Locally Differential Privacy
本文提出了一类新的本地微差分隐私下离散分布估计的隐私化方案,显著提升了中等隐私水平(其中 $1 \ll e^\epsilon \ll k$)下的性能。与现有方法相比,该方案在 $\ell_2^2$ 指标下将期望估计损失降低了 50%,在 $\ell_1$ 指标下降低了 30%,并通过紧致的下界证明了其阶次最优性。
We consider the minimax estimation problem of a discrete distribution with support size $k$ under privacy constraints. A privatization scheme is applied to each raw sample independently, and we need to estimate the distribution of the raw samples from the privatized samples. A positive number $ε$ measures the privacy level of a privatization scheme. For a given $ε,$ we consider the problem of constructing optimal privatization schemes with $ε$-privacy level, i.e., schemes that minimize the expected estimation loss for the worst-case distribution. Two schemes in the literature provide order optimal performance in the high privacy regime where $ε$ is very close to $0,$ and in the low privacy regime where $e^ε\approx k,$ respectively. In this paper, we propose a new family of schemes which substantially improve the performance of the existing schemes in the medium privacy regime when $1\ll e^ε \ll k.$ More concretely, we prove that when $3.8 < ε
研究动机与目标
- 为解决现有方案在中等隐私水平下效率不足的问题,填补该领域高效隐私化方案的空白。
- 设计一类新的隐私化方案家族,以最小化在 $\epsilon$-本地微差分隐私约束下的期望估计损失。
- 通过紧致的下界证明,所提出的方案在中等到高隐私水平($e^\epsilon \ll k$)下为阶次最优。
- 证明最优方案可被限制在输出概率比仅为 $1$ 或 $e^\epsilon$ 的极端配置中。
- 证明所提出的方案在中等隐私水平下优于 $k$-RAPPOR 和 $k$-RR,尤其在 $\ell_2^2$ 和 $\ell_1$ 估计损失方面表现更优。
提出的方法
- 作者提出了一类基于极端配置的新隐私化方案家族,其中不同输入的输出概率比为 $1$ 或 $e^\epsilon$,从而确保 $\epsilon$-本地微差分隐私。
- 通过使用有限输出字母表和结构化划分来构建方案,以在隐私约束下优化估计精度。
- 论文推导了在 $e^\epsilon \ll k$ 区域的最小最大估计损失的紧致下界,证明了所提方案在该区域的阶次最优性。
- 为隐私化样本推导出一种经验估计器,从而实现从隐私化数据中准确重构分布。
- 分析利用了概率不等式和可测集的代数构造,以建立集中性和估计误差的界。
- 通过与 $k$-RAPPOR 和 $k$-RR 的比较,验证了该方法在多种指标下显著降低了估计损失。
实验结果
研究问题
- RQ1能否设计一种新型隐私化方案,在中等隐私水平($1 \ll e^\epsilon \ll k$)下优于现有方案?
- RQ2在中等到高隐私水平($e^\epsilon \ll k$)下,估计损失的根本极限(下界)是什么?
- RQ3所提出的方案是否为阶次最优,即其与 $k$ 和 $\epsilon$ 的缩放关系是否达到最佳?
- RQ4最优隐私化方案是否可被限制在输出概率比仅为 $1$ 或 $e^\epsilon$ 的极端配置中?
- RQ5在 $\ell_2^2$ 和 $\ell_1$ 估计损失方面,新方案相较于 $k$-RAPPOR 和 $k$-RR 的改进程度如何?
主要发现
- 在中等隐私水平下,当 $3.8 < \epsilon < \ln(k/9)$ 时,所提方案将期望 $\ell_2^2$ 估计损失相比现有方案降低了 50%。
- 在相同条件下,所提方案将期望 $\ell_1$ 估计损失降低了 30%,显著优于 $k$-RAPPOR 和 $k$-RR。
- 在中等隐私水平下,所提方案在 $\ell_2^2$ 指标下相比 $k$-RR 实现了 $\Theta(k / e^\epsilon)$ 的改进因子。
- 在 $e^\epsilon \ll k$ 区域内,证明了紧致的下界,表明所提方案在该区域为阶次最优。
- 最优隐私化方案可被限制在极端配置中,即不同输入的输出概率比仅为 $1$ 或 $e^\epsilon$,从而简化了搜索空间。
- 结果表明,所提方案不仅在实验上表现更优,而且在最小最大估计风险方面也具有理论最优性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。