Skip to main content
QUICK REVIEW

[論文レビュー] Simultaneous Discrimination Prevention and Privacy Protection in Data Publishing and Mining

Sara Hajian|arXiv (Cornell University)|Jun 28, 2013
Privacy-Preserving Technologies in Data被引用数 5
ひとこと要約

本論文は、データ公開およびマイニングにおける差別の防止とプライバシー保護を同時に実現する統合フレームワークを提案する。ルール保護やフルドメイン一般化といった変換技術を導入することで、直接的および間接的差別のを排除し、k-匿名性を確保しながら、プライバシー専用ソリューションと比較して最小限のパフォーマンスオーバーヘッドでデータの有用性を維持する。

ABSTRACT

Data mining is an increasingly important technology for extracting useful knowledge hidden in large collections of data. There are, however, negative social perceptions about data mining, among which potential privacy violation and potential discrimination. Automated data collection and data mining techniques such as classification have paved the way to making automated decisions, like loan granting/denial, insurance premium computation. If the training datasets are biased in what regards discriminatory attributes like gender, race, religion, discriminatory decisions may ensue. In the first part of this thesis, we tackle discrimination prevention in data mining and propose new techniques applicable for direct or indirect discrimination prevention individually or both at the same time. We discuss how to clean training datasets and outsourced datasets in such a way that direct and/or indirect discriminatory decision rules are converted to legitimate (non-discriminatory) classification rules. In the second part of this thesis, we argue that privacy and discrimination risks should be tackled together. We explore the relationship between privacy preserving data mining and discrimination prevention in data mining to design holistic approaches capable of addressing both threats simultaneously during the knowledge discovery process. As part of this effort, we have investigated for the first time the problem of discrimination and privacy aware frequent pattern discovery, i.e. the sanitization of the collection of patterns mined from a transaction database in such a way that neither privacy-violating nor discriminatory inferences can be inferred on the released patterns. Moreover, we investigate the problem of discrimination and privacy aware data publishing, i.e. transforming the data, instead of patterns, in order to simultaneously fulfill privacy preservation and discrimination prevention.

研究の動機と目的

  • データマイニングおよび公開におけるプライバシー侵害と差別的結果という二重の脅威に対処する。
  • トレーニングデータセットおよび公開データにおける直接的および間接的差別のを検出・排除する手法を開発する。
  • プライバシー保護技術(例:k-匿名性)と差別認識変換を統合し、包括的な保護を実現する。
  • 導入された保護が、データの有用性やパフォーマンスを著しく低下させないことを保証する。
  • 規制上のプライバシーや差別禁止基準と整合する技術的対策を整備することで、法的・倫理的適合性を実現する。

提案手法

  • 直接的ルール保護(DRP)およびルール一般化(RG)を提案し、分類ルールを変換することで直接的差別を除去する。
  • 間接的ルール保護(IRP)を導入し、非差別的属性と保護対象グループとの相関関係を検出し、緩和する。
  • 直接的および間接的差別のを同時に防止するための統合変換パイプラインを開発する。
  • k-匿名性の原則を応用してα保護を構築し、一般化がプライバシー保護と差別回避の両方を満たすようにする。
  • α保護モデルをIncognitoアルゴリズムと統合し、k-匿名かつ差別的保護が施されたデータリリースを生成する。
  • 微分プライバシーとフルドメイン一般化を頻度パターン抽出に適用し、公開されたパターンからプライバシー違反や差別的推論がなされないことを保証する。

実験結果

リサーチクエスチョン

  • RQ1直接的および間接的差別を、データ有用性の低下を伴わず同時に検出・緩和することは可能か?
  • RQ2k-匿名性に基づく一般化技術は、公開データセットにおける差別の防止にも拡張可能か?
  • RQ3プライバシー保護データ公開技術は、プライバシー漏洩と差別的推論の両方に対する保護を同時に強化可能か?
  • RQ4単独のプライバシー保護と比較して、統合されたプライバシーおよび差別防止変換を適用した際のパフォーマンスおよび有用性コストはどの程度か?
  • RQ5公開出力において頻度パターンマイニングをどのようにしてプライバシー認識型および差別認識型にできるか?

主な発見

  • 提案されたα保護メカニズムは、k-匿名性と差別防止を効果的に統合しており、標準的なk-匿名化と同等のデータ歪みを示した。
  • α保護が施されたk-匿名フルドメイン一般化のサブセットは、標準k-匿名化とほぼ同等のデータ有用性を達成しており、オーバーヘッドが小さいことを示している。
  • Incognitoアルゴリズムのα保護版は、プライバシー保護と差別的でないデータリリースを生成し、実行時間は元のIncognitoとほぼ同等であった。
  • 頻度パターン抽出におけるプライバシーおよび差別防止技術の統合により、公開されたパターンからプライバシー違反や差別的推論がなされないことが保証された。
  • 実験の結果、同時的なプライバシー保護と差別防止がもたらすデータ品質への影響は最小限であり、プライバシー専用保護よりもわずかに高いにとどまった。
  • 法的根拠に基づく差別防止措置をデータ変換に組み込むことで、フレームワークは法的適合性を支援し、倫理的データ公開を可能にした。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。