Skip to main content
QUICK REVIEW

[論文レビュー] How Wrong Am I? - Studying Adversarial Examples and their Impact on Uncertainty in Gaussian Process Machine Learning Models

Kathrin Grosse, David Pfaff|arXiv (Cornell University)|Nov 17, 2017
Adversarial Robustness in Machine Learning参考文献 39被引用数 6
ひとこと要約

本稿はベイジアン不確実性の観点からガウス過程モデルにおける adversarial examples を調査し、最先端の攻撃から生じる摂動が予測不確実性を顕著に増加させることを示している。最適化されていない不確実性の閾値ですら、多くの adversarial examples を拒否できるが、攻撃を変更することでその堅牢性は覆される可能性がある。

ABSTRACT

Machine learning models are vulnerable to Adversarial Examples: minor perturbations to input samples intended to deliberately cause misclassification. Current defenses against adversarial examples, especially for Deep Neural Networks (DNN), are primarily derived from empirical developments, and their security guarantees are often only justified retroactively. Many defenses therefore rely on hidden assumptions that are subsequently subverted by increasingly elaborate attacks. This is not surprising: deep learning notoriously lacks a comprehensive mathematical framework to provide meaningful guarantees. In this paper, we leverage Gaussian Processes to investigate adversarial examples in the framework of Bayesian inference. Across different models and datasets, we find deviating levels of uncertainty reflect the perturbation introduced to benign samples by state-of-the-art attacks, including novel white-box attacks on Gaussian Processes. Our experiments demonstrate that even unoptimized uncertainty thresholds already reject adversarial examples in many scenarios. Comment: Thresholds can be broken in a modified attack, which was done in arXiv:1812.02606 (The limitations of model uncertainty in adversarial settings).

研究の動機と目的

  • ガウス過程モデルにおける adversarial perturbations が予測不確実性に与える影響を分析すること。
  • 不確実性推定が adversarial examples に対する信頼できる防御メカニズムとして機能するかどうかを評価すること。
  • ガウス過程モデルを特に標的とした新規のホワイトボックス攻撃を構築およびテストすること。
  • より洗練された攻撃に対して不確実性に基づく拒否閾値の堅牢性を評価すること。

提案手法

  • 予測における不確実性をモデル化するためにガウス過程をベイジアンフレームワークとして活用する。
  • 予測分散を用いて adversarial perturbations に対する不確実性の変化を定量化する。
  • GP モデルに対して最先端のホワイトボックス攻撃を適用し、摂動を加えた入力を生成する。
  • 高い不確実性を示す入力を不確実性閾値によって拒否する手法を採用し、adversarial examples の検出効果を評価する。
  • 不確実性に基づく検出を回避するための攻撃戦略を変更し、このような防御の脆さを示す。
  • 一般化性を確保するため、複数のデータセットおよび GP モデル設定で性能を評価する。

実験結果

リサーチクエスチョン

  • RQ1adversarial perturbations はガウス過程モデルにおいて予測不確実性をどの程度増加させるか?
  • RQ2最適化されていない不確実性閾値は、GP モデルにおける adversarial examples を効果的に検出し拒否できるか?
  • RQ3GP モデルを標的とした新規ホワイトボックス攻撃は、不確実性に基づく検出を回避する能力においてどのように比較されるか?
  • RQ4変更された攻撃は不確実性に基づく防御を回避可能か?もしそうであるなら、その方法は何か?
  • RQ5adversarial 環境において不確実性を防御メカニズムとして用いる際の限界は何か?

主な発見

  • adversarial perturbations は、異なるデータセットおよびアーキテクチャにおいても、ガウス過程モデルの予測不確実性を一貫して増加させる。
  • 最適化されていない不確実性閾値ですら、モデルの再訓練を要せず、顕著な割合の adversarial examples を拒否できる。
  • 提案された GP モデル向けホワイトボックス攻撃は、不確実性推定を巧みに利用した adversarial examples を効果的に生成する。
  • 変更された攻撃戦略により、不確実性に基づく拒否が回避可能であることが示され、このような防御が適応的攻撃に対して脆弱であることが明らかになった。
  • 先行研究(arXiv:1812.02606)で示されたように、不確実性に基づく防御は適応的攻撃に対して脆弱であることが結果から確認され、単独でモデルの不確実性に依存する堅牢性には根本的な限界があることが強調された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。