Skip to main content
QUICK REVIEW

[論文レビュー] Fingerprinting Codes and the Price of Approximate Differential Privacy

Mark Bun, Jonathan Ullman|arXiv (Cornell University)|Nov 13, 2013
Privacy-Preserving Technologies in Data参考文献 23被引用数 4
ひとこと要約

本稿は、$(\varepsilon,\delta)$-微分プライバシーを満たすアルゴリズムが多数のカウンティングクエリを正確に回答するために必要なサンプル複雑性の、タイトな情報理論的下界を確立する。近似的な微分プライバシーのコストが純粋な正確性要件よりも漸近的に大きいことを示し、下界$\tilde{\Omega}\left(\frac{\sqrt{d}\log|\mathcal{Q}|}{\alpha^2\varepsilon}\right)$を証明する。これは、対数要因を除いて既知の最良の上界と一致する。

ABSTRACT

We show new lower bounds on the sample complexity of $(\varepsilon, δ)$-differentially private algorithms that accurately answer large sets of counting queries. A counting query on a database $D \in (\{0,1\}^d)^n$ has the form "What fraction of the individual records in the database satisfy the property $q$?" We show that in order to answer an arbitrary set $\mathcal{Q}$ of $\gg nd$ counting queries on $D$ to within error $\pm α$ it is necessary that $$ n \geq ildeΩ\Bigg(\frac{\sqrt{d} \log |\mathcal{Q}|}{α^2 \varepsilon} \Bigg). $$ This bound is optimal up to poly-logarithmic factors, as demonstrated by the Private Multiplicative Weights algorithm (Hardt and Rothblum, FOCS'10). In particular, our lower bound is the first to show that the sample complexity required for accuracy and $(\varepsilon, δ)$-differential privacy is asymptotically larger than what is required merely for accuracy, which is $O(\log |\mathcal{Q}| / α^2)$. In addition, we show that our lower bound holds for the specific case of $k$-way marginal queries (where $|\mathcal{Q}| = 2^k \binom{d}{k}$) when $α$ is not too small compared to $d$ (e.g. when $α$ is any fixed constant). Our results rely on the existence of short \emph{fingerprinting codes} (Boneh and Shaw, CRYPTO'95, Tardos, STOC'03), which we show are closely connected to the sample complexity of differentially private data release. We also give a new method for combining certain types of sample complexity lower bounds into stronger lower bounds.

研究の動機と目的

  • $(\varepsilon,\delta)$-微分プライバシーを満たすアルゴリズムが、多数のカウンティングクエリを正確に回答するために必要な最小サンプル複雑性を特定すること。
  • 微分プライバシーのコストが、正確性のみを要求する場合のサンプル複雑性よりも漸近的に大きいことを確立すること。
  • 指紋コードが、プライベートデータ公開のサンプル複雑性と本質的に関連していることを示すこと。
  • 近似的な微分プライバシー下での$k$-ウェイマージナルクエリおよび一般のカウンティングクエリファミリーに対するタイトな下界を証明すること。

提案手法

  • 特定の耐性特性を持つ短い指紋コードの存在に、プライベートクエリ公開問題を還元する。
  • 異なるクエリファミリーからの下界を組み合わせる新しい合成定理を用いて、より強い全体的な下界を構築する。
  • 弱い耐性を持つ指紋コード(特にTardosコード)を用い、確率的および集中度解析を用いてその強い耐性を証明する。
  • 完全耐性への還元を指紋コードの弱い耐性から行い、サンプル複雑性の下界を確立する。
  • Markovの不等式と二項分布および三角関数積分の尾部確率を制御するための不等式を用い、解析における誤差確率を制御する。
  • 指紋コードとクエリファミリーのVC次元との関係を活用し、情報理論的限界を導出する。

実験結果

リサーチクエスチョン

  • RQ1$(\varepsilon,\delta)$-微分プライバシーを満たすアルゴリズムが、$|\mathcal{Q}|$個のカウンティングクエリを誤差$\pm\alpha$で正確に回答するための最小サンプルサイズは何か?
  • RQ2$(\varepsilon,\delta)$-微分プライバシーの要件は、正確な統計的推定に必要なサンプル複雑性を超えて、漸近的なコストを負うか?
  • RQ3指紋コードは、微分プライベートデータ公開のサンプル複雑性とどのように関連しているか?
  • RQ4$k$-ウェイマージナルなどのクエリファミリーの構造的性質を用いて、サンプル複雑性の下界を改善またはタイトにすることができるか?
  • RQ5Private Multiplicative Weightsアルゴリズムのサンプル複雑性は、対数要因を除いて最適か?

主な発見

  • $(\varepsilon,\delta)$-微分プライバシーを満たす$|\mathcal{Q}|$個のカウンティングクエリの公開に必要なサンプル複雑性は、少なくとも$\tilde{\Omega}\left(\frac{\sqrt{d}\log|\mathcal{Q}|}{\alpha^2\varepsilon}\right)$であり、これは多対数的要因を除いてタイトである。
  • この下界は、プライバシーなしで正確な推定に必要な$O(\log|\mathcal{Q}|/\alpha^2)$のサンプル複雑性よりも厳密に大きいことを示し、微分プライバシーの非自明なコストを確立する。
  • $k$-ウェイマージナルクエリで$|\mathcal{Q}| = 2^k \binom{d}{k}$の場合、$\alpha$が定数であれば下界は$\Omega(\sqrt{d}/\alpha^2\varepsilon)$であり、$\alpha$が小さい場合には$\Omega(k/\varepsilon)$である。
  • Tardosの指紋コードが弱く耐性を持つことが示され、この性質が完全耐性への還元を経て強い下界を導出するのに用いられる。
  • 解析により、確率的に生成されたTardosコードブックに$\Omega(n^{3/2}\log(n/\xi))$個のマークされた列が含まれることが示され、これは下界構築にとって不可欠である。
  • 本稿は、Private Multiplicative Weightsアルゴリズムが、対数要因を除いて最適なサンプル複雑性を達成することを確立し、下界のタイトさを確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。