Skip to main content
QUICK REVIEW

[論文レビュー] On the Complexity of $t$-Closeness Anonymization and Related Problems

Hongyu Liang, Hao Yuan|arXiv (Cornell University)|Jan 9, 2013
Privacy-Preserving Technologies in Data参考文献 27被引用数 4
ひとこと要約

本稿は、$t$-closeness匿名化の最初の体系的理論的分析を提示し、任意の定数 $t \in [0,1)$ に対して最適な $t$-closeness一般化を見つけることがNP困難であることを証明する。さらに、正確かつ固定パラメータ付きのアルゴリズムを提供し、$k$-匿名性および $l$-多様性に関する未解決の複雑性の問題を解消し、2-多様性が多項式時間で解けることを確立する。

ABSTRACT

An important issue in releasing individual data is to protect the sensitive information from being leaked and maliciously utilized. Famous privacy preserving principles that aim to ensure both data privacy and data integrity, such as $k$-anonymity and $l$-diversity, have been extensively studied both theoretically and empirically. Nonetheless, these widely-adopted principles are still insufficient to prevent attribute disclosure if the attacker has partial knowledge about the overall sensitive data distribution. The $t$-closeness principle has been proposed to fix this, which also has the benefit of supporting numerical sensitive attributes. However, in contrast to $k$-anonymity and $l$-diversity, the theoretical aspect of $t$-closeness has not been well investigated. We initiate the first systematic theoretical study on the $t$-closeness principle under the commonly-used attribute suppression model. We prove that for every constant $t$ such that $0\leq t<1$, it is NP-hard to find an optimal $t$-closeness generalization of a given table. The proof consists of several reductions each of which works for different values of $t$, which together cover the full range. To complement this negative result, we also provide exact and fixed-parameter algorithms. Finally, we answer some open questions regarding the complexity of $k$-anonymity and $l$-diversity left in the literature.

研究の動機と目的

  • 属性削除フレームワーク下での $t$-closenessプライバシー・モデルの体系的理論的研究を開始すること。
  • 任意の定数 $t \in [0,1)$ に対して最適な $t$-closeness一般化を見つける計算複雑性を特定すること。
  • $t$-closeness問題を解くための正確および固定パラメータ付きのアルゴリズムを開発すること。
  • 先行文献で未解決のまま残っていた $k$-匿名性および $l$-多様性に関する複雑性の未解決問題を解消すること。
  • 2-多様性のための最初の多項式時間アルゴリズムを確立し、その計算的 tractability を決定すること。

提案手法

  • 異なる $t$ 値の範囲に合わせた複数の還元を用いて、$t$-closeness の NP 困難性を証明し、区間 $[0,1)$ 全体をカバーする。
  • 入力テーブルからハイパーグラフを構築し、2-多様性を単体条件を満たす単体マッチング問題に還元する。
  • 既知の多項式時間アルゴリズムを活用して、単体マッチング問題を解き、2-多様性を多項式時間で解く。
  • 動的計画法およびパラメータ化複雑性技術を用いて、$t$-closeness のための正確および固定パラメータ付きのアルゴリズムを設計する。
  • ケース解析および帰納的推論を用いて、最適な2-多様な分割が、異なる感応属性値を持つサイズ2または3のグループに制限できることを証明する。
  • 構築されたハイパーグラフが単体条件を満たすことを保証し、既存の多項式時間マッチングアルゴリズムの適用を可能にする。

実験結果

リサーチクエスチョン

  • RQ1任意の定数 $t \in [0,1)$ に対して、最適な $t$-closeness一般化を見つける問題は NP 困難であるか?
  • RQ2$t$-closeness 問題は、正確または固定パラメータ付きのアルゴリズムを用いて効率的に解けるか?
  • RQ32-多様性は多項式時間で解けるか? もしそうなら、どのような構造的条件下で?
  • RQ4特に $k$ や $l$ の小さい値に対して、属性削除モデル下での $k$-匿名性および $l$-多様性の計算複雑性は何か?
  • RQ5$t$-closeness に対して、保証可能な近似比を持つ多項式時間近似アルゴリズムを設計できるか?

主な発見

  • $t$-closeness 問題は、$[0,1)$ の任意の定数 $t$ に対して、複数の特化された還元を用いて証明された。
  • 小規模または構造的特徴を持つインスタンスに対して有効な正確アルゴリズムおよび固定パラメータ付きアルゴリズムが開発された。
  • 単体条件を満たす単体マッチング問題への還元により、2-多様性のための最初の多項式時間アルゴリズムが確立された。
  • 最適な2-多様な分割は、異なる感応属性値を持つサイズ2または3のグループに制限でき、計算が効率的に行える。
  • $k$-匿名性の複雑性は条件付きで改善され、新たな理論的知見により先行研究で未解決のままだった問題が解消された。
  • 著者らは、$t$-closeness の最良近似比が $O(\log n)$ である可能性を仮説し、将来的な近似アルゴリズムの方向性を示唆している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。