Skip to main content
QUICK REVIEW

[論文レビュー] SEC4SR: A Security Analysis Platform for Speaker Recognition

Guangke Chen, Zhe Zhao|arXiv (Cornell University)|Sep 4, 2021
Adversarial Robustness in Machine Learning参考文献 95被引用数 6
ひとこと要約

SEC4SRは、話者認識システムにおける包括的で体系的なセキュリティ分析プラットフォームであり、15の攻撃(適応的バリエーションを含む)および23の防御(新規の特徴レベル変換を含む)を、4つの認識タスクおよび5つのデータセットを対象として体系的に評価可能である。本研究では、特徴レベル変換と adversarial training を組み合わせることで、あらゆる攻撃タイプに対して最も強力な防御が得られることを示している。

ABSTRACT

Adversarial attacks have been expanded to speaker recognition (SR). However, existing attacks are often assessed using different SR models, recognition tasks and datasets, and only few adversarial defenses borrowed from computer vision are considered. Yet,these defenses have not been thoroughly evaluated against adaptive attacks. Thus, there is still a lack of quantitative understanding about the strengths and limitations of adversarial attacks and defenses. More effective defenses are also required for securing SR systems. To bridge this gap, we present SEC4SR, the first platform enabling researchers to systematically and comprehensively evaluate adversarial attacks and defenses in SR. SEC4SR incorporates 4 white-box and 2 black-box attacks, 24 defenses including our novel feature-level transformations. It also contains techniques for mounting adaptive attacks. Using SEC4SR, we conduct thus far the largest-scale empirical study on adversarial attacks and defenses in SR, involving 23 defenses, 15 attacks and 4 attack settings. Our study provides lots of useful findings that may advance future research: such as (1) all the transformations slightly degrade accuracy on benign examples and their effectiveness vary with attacks; (2) most transformations become less effective under adaptive attacks, but some transformations become more effective; (3) few transformations combined with adversarial training yield stronger defenses over some but not all attacks, while our feature-level transformation combined with adversarial training yields the strongest defense over all the attacks. Extensive experiments demonstrate capabilities and advantages of SEC4SR which can benefit future research in SR.

研究の動機と目的

  • 話者認識(SR)における敵対的攻撃および防御の体系的かつ包括的な評価プラットフォームの不足に応えること。
  • 多様なSRモデル、データセット、防御手法をカバーする非適応的および適応的攻撃を統合的に評価可能な統一フレームワークを提供すること。
  • 入力レベルおよび特徴レベル変換の有効性、特に適応的攻撃状況下での有効性を調査すること。
  • adversarial training とさまざまな防御技術の相乗効果を評価し、最も強固な防御構成を同定すること。

提案手法

  • BPDA、EOT、NESを用いて防御回避を可能にする適応的バージョンに拡張された、4つのホワイトボックスおよび2つのブラックボックス敵対的攻撃を統合。
  • 24の防御メカニズムを実装。22の入力および特徴レベル変換(時間領域・周波数領域、音声圧縮、および新規の特徴レベル操作)と2つの adversarial training 方法を含む。
  • 3つの主流のSRSと5つの標準化された音声データセットをサポート。4つの認識タスク(例:CSI-E、SV、OSI、CSI-NE)をカバー。
  • 画像ベースの指標だけでなく、人間の聴覚認識に整合する歪み指標を導入。
  • 攻撃者知識の増加に伴う防御の堅牢性を評価可能な、非適応的および適応的攻撃設定の両方を設定可能に。
  • 攻撃成功の定量的評価のための8つの攻撃指標と、防御効果の定量的評価のための3つの防御指標を提供。

実験結果

リサーチクエスチョン

  • RQ1異なる攻撃設定およびモデル下で、SRSはどの程度敵対的攻撃に対して脆弱であるか?
  • RQ2非適応的攻撃に対して、入力および特徴レベル変換はどの程度有効か?
  • RQ3攻撃者が防御の詳細を完全に把握している状況(適応的攻撃)下でも、これらの変換は有効性を保つのか?
  • RQ4adversarial training と特徴レベル変換を組み合わせることで、より強力で汎用性の高い防御が得られるか?

主な発見

  • すべての入力および特徴レベル変換は、クリーンな正確性をわずかに低下させるが、攻撃タイプによってその有効性は顕著に異なる。
  • 大多数の変換は適応的攻撃下で有効性を失うが、特に特徴レベルの変換はこのような状況下でより有効になることがある。
  • adversarial training と特定の変換を組み合わせることで防御の堅牢性が向上するが、これは特定の攻撃タイプに限定される。
  • 提案された特徴レベル変換は、adversarial training と組み合わせることで、15の攻撃設定すべてにおいて最も強力な防御を達成し、他の防御組み合わせを上回る。
  • adversarial training だけでは適応的攻撃に対して不十分であることが示され、ハイブリッドな防御戦略の必要性が浮き彫りになった。
  • コンピュータビジョン分野の既存防御(例:画像圧縮)は、音声信号のドメイン固有の特性のため、SRでは効果を示さない。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。