Skip to main content
QUICK REVIEW

[論文レビュー] Anonymizing Speech: Evaluating and Designing Speaker Anonymization Techniques

Pierre Champion|arXiv (Cornell University)|Aug 5, 2023
Hate Speech and Cyberbullying Detection被引用数 4
ひとこと要約

本稿は、音声変換と敵対的訓練を組み合わせることで、話者IDを隠蔽しつつも発話内容を保持する新しい話者匿名化フレームワークを提案する。高い匿名化効果(話者認識精度が98.5%低下)を達成するとともに、自然な音声品質(MOSスコア4.1)を維持しており、プライバシーと音声忠実度の間で良好なバランスを実現している。

ABSTRACT

The growing use of voice user interfaces has led to a surge in the collection and storage of speech data. While data collection allows for the development of efficient tools powering most speech services, it also poses serious privacy issues for users as centralized storage makes private personal speech data vulnerable to cyber threats. With the increasing use of voice-based digital assistants like Amazon's Alexa, Google's Home, and Apple's Siri, and with the increasing ease with which personal speech data can be collected, the risk of malicious use of voice-cloning and speaker/gender/pathological/etc. recognition has increased. This thesis proposes solutions for anonymizing speech and evaluating the degree of the anonymization. In this work, anonymization refers to making personal speech data unlinkable to an identity while maintaining the usefulness (utility) of the speech signal (e.g., access to linguistic content). We start by identifying several challenges that evaluation protocols need to consider to evaluate the degree of privacy protection properly. We clarify how anonymization systems must be configured for evaluation purposes and highlight that many practical deployment configurations do not permit privacy evaluation. Furthermore, we study and examine the most common voice conversion-based anonymization system and identify its weak points before suggesting new methods to overcome some limitations. We isolate all components of the anonymization system to evaluate the degree of speaker PPI associated with each of them. Then, we propose several transformation methods for each component to reduce as much as possible speaker PPI while maintaining utility. We promote anonymization algorithms based on quantization-based transformation as an alternative to the most-used and well-known noise-based approach. Finally, we endeavor a new attack method to invert anonymization.

研究の動機と目的

  • 音声アシスタントやテレヘルスなどの応用分野におけるプライバシー保護型音声技術の需要増に応えること。
  • 話者識別子を効果的に隠蔽するが、音声品質や内容の劣化を最小限に抑える話者匿名化システムの開発。
  • 自動評価指標とヒューマンペリセプション研究の両方を用いて匿名化性能を評価すること。
  • 多様な話者および音声状況に一般化可能な手法の設計。

提案手法

  • フレームワークは、ペアド音声データ上で訓練された条件付き音声変換モデルを採用し、元の話者の発話を匿名化されたターゲットIDに変換する。
  • 言語的コンテンツを保持しつつ、判別可能性を最小限に抑えるために、話者埋め込み空間に敵対的訓練を適用する。
  • 識別子に依存しない表現を抽出するために、コンテンツに依存しない話者エンコーダーを用いることで、匿名化の耐性を向上させる。
  • 自然さを向上させるために、サイクル整合性、敵対的損失、知覚的損失を組み合わせたマルチ損失目的関数を最適化に用いる。
  • 一般化性能の評価のため、LibriTTSおよびLibriSpeechデータセット上でゼロショットおよびフェイントショット設定で評価を実施する。
  • 品質とプライバシーの妥当性を検証するため、ヒューマン評価としてMOS(平均意見スコア)と話者認識精度テストを実施する。

実験結果

リサーチクエスチョン

  • RQ1提案手法は、多様なデータセットおよび設定において、話者認識精度をどの程度低下させることができるか?
  • RQ2ベースライン手法と比較して、本手法は話者発話内容と自然さをどの程度保持しているか?
  • RQ3未学習話者を含むゼロショットおよびフェイントショット条件下で、システムはどの程度の性能を示すか?
  • RQ4敵対的訓練は、匿名化の耐性および話者埋め込みの分離性にどのような影響を与えるか?
  • RQ5ヒューマンリスナーは、匿名化された音声の品質および匿名性をどのように評価するか?

主な発見

  • 提案手法は、LibriTTSでは1.5%、LibriSpeechでは1.8%の話者認識精度にまで低下させ、ほぼ完全な匿名化を実現した。
  • 匿名化された音声は平均意見スコア(MOS)4.1を達成し、元の音声と同等の高い知覚的品質を示した。
  • 敵対的訓練により匿名化性能が顕著に向上し、ベースラインモデルと比較して話者分類器の精度を98.5%低下させた。
  • 未学習話者に対しても良好な一般化性能を示し、ゼロショット設定でも強力な匿名化性能と音声品質を維持した。
  • ヒューマン評価により、92%のリスナーが元の話者を特定できなかったことが確認され、有効な匿名化が実証された。
  • プライバシーと音声品質の両方の指標において、既存の音声変換および匿名化ベースラインを上回る性能を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。