Skip to main content
QUICK REVIEW

[論文レビュー] DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake Voices

Run Wang, Felix Juefei-Xu|arXiv (Cornell University)|May 28, 2020
Music and Audio Processing参考文献 49被引用数 17
ひとこと要約

DeepSonar は、スプライターレコognition モデル内のレイヤーごとの活性化パターンを分析することで、AI 生成のなりすまし声を検出する、ニューロン行動に基づく新規アプローチを提案する。複数のデータセットで平均 98.1% の精度と低誤検出率(約 2%)を達成し、ボイスコンバージョンや実世界のノイズ攻撃に対しても高い耐性を示す。

ABSTRACT

With the recent advances in voice synthesis, AI-synthesized fake voices are indistinguishable to human ears and widely are applied to produce realistic and natural DeepFakes, exhibiting real threats to our society. However, effective and robust detectors for synthesized fake voices are still in their infancy and are not ready to fully tackle this emerging threat. In this paper, we devise a novel approach, named \emph{DeepSonar}, based on monitoring neuron behaviors of speaker recognition (SR) system, \ie, a deep neural network (DNN), to discern AI-synthesized fake voices. Layer-wise neuron behaviors provide an important insight to meticulously catch the differences among inputs, which are widely employed for building safety, robust, and interpretable DNNs. In this work, we leverage the power of layer-wise neuron activation patterns with a conjecture that they can capture the subtle differences between real and AI-synthesized fake voices, in providing a cleaner signal to classifiers than raw inputs. Experiments are conducted on three datasets (including commercial products from Google, Baidu, \etc) containing both English and Chinese languages to corroborate the high detection rates (98.1\% average accuracy) and low false alarm rates (about 2\% error rate) of DeepSonar in discerning fake voices. Furthermore, extensive experimental results also demonstrate its robustness against manipulation attacks (\eg, voice conversion and additive real-world noises). Our work further poses a new insight into adopting neuron behaviors for effective and robust AI aided multimedia fakes forensics as an inside-out approach instead of being motivated and swayed by various artifacts introduced in synthesizing fakes.

研究の動機と目的

  • 人間の声と区別がつかないほど精巧に生成される AI 生成のなりすまし声が引き起こす詐欺やフェイクニュースの脅威に対処すること。
  • 現代のシステムではしばしば存在しないか、微細な合成アーティファクトに依存する既存の検出手法の限界を克服すること。
  • 騒音環境や後処理の操作が加えられた状況を含む実世界の条件でも効果的かつ耐性を持つ検出フレームワークを開発すること。
  • 外部アーティファクトではなく、内部ニューロン行動に基づく「内側から見る」検出アプローチを検討し、一般化性能と解釈可能性を向上させること。

提案手法

  • 事前に訓練済みのスプライターレコognition(SR)モデルにおけるレイヤーごとのニューロン活性化パターンを監視し、本物の声となりすまし声を区別するための特徴表現を抽出する。
  • 音声合成時に生じる人工的アーティファクトに依存せず、DNN の内部表現を分類の信号源として利用する。
  • 複数のレイヤーにわたるニューロン活性化パターンを集約して分類器を学習させ、本物の人間の声と AI 生成のなりすまし声を区別する。
  • 多様な声質や言語(英語および中国語)にわたる一般化を向上させるために、データ拡張および正規化技術を適用する。
  • 実世界の展開状況を模倣するため、テスト入力にボイスコンバージョンおよび追加的な実世界のノイズ(例:風、雨、笑い)を適用し、耐性を評価する。
  • bispectral アーティファクト分析を用いたベースラインと比較し、操作攻撃下でも優れた耐性を示すことを実証する。

実験結果

リサーチクエスチョン

  • RQ1スプライターレコognition モデル内のレイヤーごとのニューロン活性化パターンは、本物の人間の声と AI 生成のなりすまし声を効果的に区別できるか?
  • RQ2本研究で提案するニューロン行動に基づく検出手法は、実世界のノイズ環境およびボイスコンバージョン攻撃下でどのように性能を発揮するか?
  • RQ3外部アーティファクト(例:bispectral 特徴)に依存するのではなく、DNN の内部動作を監視することで、より優れた検出性能と耐性が得られるか?
  • RQ4本手法は、さまざまな言語および商用 TTS/VC システム(例:Google、Baidu)に対しても一般化できるか?

主な発見

  • DeepSonar は、商用 TTS やボイスクラーニングシステムを含む、3 つの多様なデータセットで平均 98.1% の検出精度を達成した。
  • 誤検出率は約 2% にとどまり、本物の声となりすまし声の区別において高い信頼性を示している。
  • ボイスコンバージョン攻撃下でも、5 つのノイズ強度で平均して 7% 未満の精度低下に抑えられ、高い性能を維持した。
  • 追加の実世界ノイズでは、屋内ノイズ(例:花火、列車)では 18% 未満、風や雨では最大 25% の性能低下を示したが、SNR > 35 dB の条件下では依然として有効であった。
  • あらゆるタイプの操作攻撃において、bispectral アーティファクトベースのベースラインを顕著に上回り、一貫して高い AUC スコアを示した。
  • 特にノイズ除去処理を適用した場合、複雑なノイズ混合物に対しても耐性を示し、今後の研究において音声強化技術と統合可能である可能性を示唆している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。