Skip to main content
QUICK REVIEW

[論文レビュー] ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning Models

Ahmed Salem, Yang Zhang|arXiv (Cornell University)|Jun 4, 2018
Adversarial Robustness in Machine Learning被引用数 13
ひとこと要約

本稿では、機械学習モデルに対する、モデルおよびデータに依存しない-membership inference 攻撃の初の提案を行う。これにより、最小限の仮定のもとで、1つのシャドウモデル、あるいはまったくシャドウモデルを用いなくても、効果的な攻撃が可能であることが示された。著者らはさらに、ドロップアウトとモデルスタッキングの2つの実用的防御策を提案し、8つの多様なデータセットにおいて、モデルの有用性を保ちながら攻撃成功確率を顕著に低減した。

ABSTRACT

Machine learning (ML) has become a core component of many real-world applications and training data is a key factor that drives current progress. This huge success has led Internet companies to deploy machine learning as a service (MLaaS). Recently, the first membership inference attack has shown that extraction of information on the training set is possible in such MLaaS settings, which has severe security and privacy implications. However, the early demonstrations of the feasibility of such attacks have many assumptions on the adversary, such as using multiple so-called shadow models, knowledge of the target model structure, and having a dataset from the same distribution as the target model's training data. We relax all these key assumptions, thereby showing that such attacks are very broadly applicable at low cost and thereby pose a more severe risk than previously thought. We present the most comprehensive study so far on this emerging and developing threat using eight diverse datasets which show the viability of the proposed attacks across domains. In addition, we propose the first effective defense mechanisms against such broader class of membership inference attacks that maintain a high level of utility of the ML model.

研究の動機と目的

  • 先行研究と比較して、仮定を大幅に緩和した条件下でも、機械学習モデルに対する-membership inference 攻撃が実現可能であることを示すこと。
  • 攻撃者がターゲットモデルのアーキテクチャや訓練データの分布にアクセスできない状況でも、攻撃が有効に機能することを示すこと。
  • -membership inference のリスクを低下させるが、モデル性能を著しく低下させない、新たな防御メカニズムの開発と評価を行うこと。
  • 多様な機械学習応用およびデータセットにおいて、-membership inference の脅威が一般性と実用性を有することを確立すること。

提案手法

  • 複数のシャドウモデルを必要としない、単一のシャドウモデルを用いた攻撃を提案し、コストと複雑さを著しく低減する。
  • 異なるデータセットを用いてシャドウモデルを訓練するデータ転送攻撃を導入し、ターゲットモデルの訓練データ分布にアクセスできない状況でも攻撃を可能にする。
  • シャドウモデルを一切構築しない、非教師あり-membership inference 攻撃を提案し、ターゲットモデルの出力挙動のみに依存する。
  • 訓練時にドロップアウト正則化を適用し、モデルの過適合を軽減することで、-membership レイクレージを低減する。
  • 複数のベースモデルを階層的アンサンブルに統合するモデルスタッキングを用い、-membership inference に対してより高い耐性を実現する。
  • CNNやMLPなど、さまざまな機械学習モデルを用いて、画像、テキスト、表形式の8つの多様なデータセットで攻撃および防御のパフォーマンスを評価する。

実験結果

リサーチクエスチョン

  • RQ1複数のモデルではなく、1つのシャドウモデルのみを用いた場合でも、-membership inference 攻撃が有効に機能するか?
  • RQ2ターゲットモデルの訓練データと同じ分布からのデータにアクセスできない状況でも、-membership inference をどの程度実行できるか?
  • RQ3シャドウモデルを一切構築しない、完全に非教師ありの状況でも -membership inference が可能か?
  • RQ4ドロップアウトやモデルスタッキングといった防御機構が、-membership inference の成功確率を顕著に低下させつつ、高いモデル精度を維持できるか?

主な発見

  • 1つのシャドウモデルを用いることで、10個のシャドウモデルを用いた従来手法とほぼ同等の攻撃性能(CIFAR-100とCNNを用いた場合、適合率0.95、再現率0.95)を達成した。
  • データ転送攻撃—関連のないデータセットを用いてシャドウモデルを訓練する—は、強力な-membership inference の性能を示し、広範な適用可能性を示した。
  • シャドウモデルを一切必要としない非教師あり攻撃も、依然として効果的な-membership inference を実現しており、脅威が広範にわたることを証明した。
  • ドロップアウトを用いた防御により、すべての評価対象データセットで-membership inference 攻撃の成功確率が顕著に低下したが、高いモデル精度を維持した。
  • モデルスタッキングを防御に用いる場合も、攻撃性能を効果的に低減でき、ターゲットモデルの予測有用性への影響は最小限にとどまった。
  • 両方の防御を組み合わせることで、テストされた8つの多様なデータセットすべてで-membership inference の成功確率が著しく低下した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。