Skip to main content
QUICK REVIEW

[論文レビュー] An Overview of Backdoor Attacks Against Deep Neural Networks and Possible Defences

Wei Guo, Benedetta Tondi|arXiv (Cornell University)|Nov 16, 2021
Adversarial Robustness in Machine Learning被引用数 5
ひとこと要約

この論文は、深層ニューラルネットワーク(DNN)におけるバックドア攻撃について包括的なサーベイを提供し、攻撃者が訓練段階での制御能力と防御者が検証能力を持つという観点から分類している。攻撃タイプ、検出手法、防御策(敵対的例検出のためのバックドアベースのウォーターマーキングやトラップドア・ハニーポットを含む)をレビューし、応用シナリオに応じた強み、弱み、適性を強調している。

ABSTRACT

Together with impressive advances touching every aspect of our society, AI technology based on Deep Neural Networks (DNN) is bringing increasing security concerns. While attacks operating at test time have monopolised the initial attention of researchers, backdoor attacks, exploiting the possibility of corrupting DNN models by interfering with the training process, represents a further serious threat undermining the dependability of AI techniques. In a backdoor attack, the attacker corrupts the training data so to induce an erroneous behaviour at test time. Test time errors, however, are activated only in the presence of a triggering event corresponding to a properly crafted input sample. In this way, the corrupted network continues to work as expected for regular inputs, and the malicious behaviour occurs only when the attacker decides to activate the backdoor hidden within the network. In the last few years, backdoor attacks have been the subject of an intense research activity focusing on both the development of new classes of attacks, and the proposal of possible countermeasures. The goal of this overview paper is to review the works published until now, classifying the different types of attacks and defences proposed so far. The classification guiding the analysis is based on the amount of control that the attacker has on the training process, and the capability of the defender to verify the integrity of the data used for training, and to monitor the operations of the DNN at training and test time. As such, the proposed analysis is particularly suited to highlight the strengths and weaknesses of both attacks and defences with reference to the application scenarios they are operating in.

研究の動機と目的

  • 攻撃者による制御能力と防御者の検証能力に基づいて、DNNにおけるバックドア攻撃と防御を体系的に分類すること。
  • バックドア攻撃のための形式的脅威モデルを定式化し、完全制御と部分的制御のシナリオを区別すること。
  • 既存のバックドア検出および除去手法の有効性と限界を評価すること。
  • バックドアをDNNの知的財産権保護のための動的ウォーターマーキングと敵対的例検出に二重用途で利用できる可能性を検討すること。
  • 微調整、トランスファーラーニング、ブラックボックス攻撃に対する耐性に関する未解決の課題を特定すること。

提案手法

  • 攻撃者がモデルを訓練する完全制御(full control)と、データや勾配を操作する部分的制御(partial control)のシナリオに分類してバックドア攻撃を分類する。
  • 訓練時およびテスト時の攻撃者能力と防御者制約を分析するための形式的脅威モデルフレームワークを提案する。
  • 訓練時(例:訓練ダイナミクスのモニタリング)とテスト時(例:活性化解析、入力プローブ)の両方で動作する検出手法をレビューする。
  • バックドアベースのウォーターマーキングを検討:特定のトリガーを有する汚染された訓練データを用いてバックドアを動的ウォーターマーキングとして埋め込む。
  • トラップドア・ハニーポットの概念を導入:低エネルギーのトリガーを用いてDNNを訓練し、トリガーに近い入力を特定することで敵対的例を検出する。
  • 微調整およびトランスファーラーニング下でのウォーターマーキングおよび防御の耐性を評価し、主要な脆弱性を同定する。

実験結果

リサーチクエスチョン

  • RQ1訓練プロセスにおける攻撃者の制御レベルの違いが、バックドア攻撃の実現可能性と隠蔽性にどのように影響するか?
  • RQ2訓練時とテスト時のバックドア検出手法の有効性に、どのような主な差異があるか?
  • RQ3動的ウォーターマーキングを用いたバックドアは、DNNの知的財産権保護にどの程度有効に利用できるか?
  • RQ4トラップドアベースのハニーポットは、標的型敵対的例を効果的に検出できるか、その限界は何か?
  • RQ5なぜバックドアベースのウォーターマーキングの耐性はトランスファーラーニング下で特に弱体化するのか、そしてどのように改善できるか?

主な発見

  • 汚染によるバックドアベースのウォーターマーキングは、ASR(攻撃成功確率)を検証指標として用いることで所有権を確立可能だが、微調整後および特にトランスファーラーニング後には耐性が著しく低下する。
  • 低エネルギーのトリガーを活用することで、トラップドア・ハニーポット防御は標的型敵対的例を高い正確性で検出できる。これは、敵対的攻撃が低エネルギーのトリガーを狙いやすいという性質を利用している。
  • 既存の防御策は特定の脅威モデル下でのみ有効であり、非標的攻撃やブラックボックス攻撃に対してはしばしば失敗する。
  • トランスファーラーニングは動的ウォーターマーキングの耐性を著しく損なうため、実用的知的財産権保護の主要な制限要因である。
  • バックドアをウォーターマーキングと敵対的検出に二重用途で利用することは、セキュリティとパケット容量の面で静的ウォーターマーキングに比べて制限が大きいが、二重用途の性質を示している。
  • どの防御策も普遍的に有効ではない。耐性は攻撃シナリオ、モデルアーキテクチャ、および微調整などのトレーニング後処理に強く依存する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。