[論文レビュー] The Hidden Vulnerability of Watermarking for Deep Neural Networks.
本論文は、水増しパターンを隠すための前処理関数と、分布外データにおけるファインチューニング戦略を用いることで、最先端のDNN水増し技術を回避する、知識を要しない新しい水増し除去攻撃を提案する。この攻撃は、水増し手法やトレーニングサンプルに関する事前知識がなくても、水増しの除去に非常に高い成功率を達成する。
Watermarking has shown its effectiveness in protecting the intellectual property of Deep Neural Networks (DNNs). Existing techniques usually embed a set of carefully-crafted sample-label pairs into the target model during the training process. Then ownership verification is performed by querying a suspicious model with those watermark samples and checking the prediction results. These watermarking solutions claim to be robustness against model transformations, which is challenged by this paper. We design a novel watermark removal attack, which can defeat state-of-the-art solutions without any prior knowledge of the adopted watermarking technique and training samples. We make two contributions in the design of this attack. First, we propose a novel preprocessing function, which embeds imperceptible patterns and performs spatial-level transformations over the input. This function can make the watermark sample unrecognizable by the watermarked model, while still maintaining the correct prediction results of normal samples. Second, we introduce a fine-tuning strategy using unlabelled and out-of-distribution samples, which can improve the model usability in an efficient manner. Extensive experimental results indicate that our proposed attack can effectively bypass existing watermarking solutions with very high success rates.
研究の動機と目的
- 既存のDNN水増し技術がモデル変換に対して主張する耐性を検証すること。
- 水増し手法やトレーニングサンプルに関する事前知識が不要な水増し除去攻撃を設計すること。
- 通常の入力に対してモデルの精度を維持しつつ、水増しサンプルを水増し済みモデルが認識できなくなるようにすること。
- 水増し除去後にモデルの有用性を効率的なファインチューニングにより向上させること。
提案手法
- 入力に知覚できないパターンを埋め込み、水増し信号をぼかすための空間的変換を施す、新たな前処理関数を導入する。
- 前処理関数は、通常のサンプルに対して正しい予測を保証するとともに、水増しサンプルを水増し済みモデルが認識できなくする。
- ラベルなしの分布外サンプルを用いたファインチューニング戦略を適用し、水増し除去後のモデル性能を回復する。
- 元の水増しトレーニングデータや水増し手法に関する知識が不要である。
- 一般化能力を示すために、複数のDNNアーキテクチャと水増しスキームで評価を行う。
- 攻撃ワークフローは、水増し入力を前処理し、モデル適合による水増しを除去し、モデルの有用性を維持するためのファインチューニングを含む。
実験結果
リサーチクエスチョン
- RQ1水増し手法やトレーニングサンプルに関する事前知識がなくても、水増し除去攻撃を設計できるか?
- RQ2知覚できない前処理関数は、モデルの精度を維持しつつ、どの程度水増しパターンを隠すことができるか?
- RQ3分布外データにおけるファインチューニングは、水増し除去後のモデルの有用性回復にどの程度効果的か?
- RQ4提案された攻撃は、多様なDNNモデルで、最先端の水増しスキームを高い成功率で回避できるか?
主な発見
- 提案された攻撃は、最先端のDNN水増しスキームからの水増し除去において非常に高い成功率を達成する。
- 水増し手法やトレーニングサンプルに関する知識がなくても、攻撃は効果的に水増しを回避する。
- 前処理関数により、水増し済みモデルが水増しサンプルを認識できなくなる一方で、通常の入力に対しては正しい予測が維持される。
- ラベルなしの分布外サンプルを用いたファインチューニングは、水増し除去後のモデルの有用性を効果的に回復する。
- 広範な実験により、攻撃の有効性が複数のDNNアーキテクチャと水増し技術で確認された。
- 攻撃は、さまざまな水増し防御に対して強い一般化性と耐性を示した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。