Skip to main content
QUICK REVIEW

[논문 리뷰] Defending Your Voice: Adversarial Attack on Voice Conversion

Chien‐Yu Huang, Yist Y. Lin|arXiv (Cornell University)|2020. 05. 18.
Speech Recognition and Synthesis참고 문헌 41인용 수 5
한 줄 요약

이 논문은 개인의 화자 신원을 보호하기 위해 음성 변환 시스템에 대한 최초의 적대적 공격을 제시한다. 인간이 인지하지 못하는 노이즈를 화자 발화에 주입함으로써, 이 방법은 제로샷 음성 변환 모델을 교란시켜 변환된 출력이 방어된 화자와 더 이상 유사하지 않게 만들지만, 적대적 입력은 화이트박스 및_BLK박스 설정에서 모두 객관적이고 주관적인 평가에서 진짜 음성과 구분되지 않는다.

ABSTRACT

Substantial improvements have been achieved in recent years in voice conversion, which converts the speaker characteristics of an utterance into those of another speaker without changing the linguistic content of the utterance. Nonetheless, the improved conversion technologies also led to concerns about privacy and authentication. It thus becomes highly desired to be able to prevent one's voice from being improperly utilized with such voice conversion technologies. This is why we report in this paper the first known attempt to perform adversarial attack on voice conversion. We introduce human imperceptible noise into the utterances of a speaker whose voice is to be defended. Given these adversarial examples, voice conversion models cannot convert other utterances so as to sound like being produced by the defended speaker. Preliminary experiments were conducted on two currently state-of-the-art zero-shot voice conversion models. Objective and subjective evaluation results in both white-box and black-box scenarios are reported. It was shown that the speaker characteristics of the converted utterances were made obviously different from those of the defended speaker, while the adversarial examples of the defended speaker are not distinguishable from the authentic utterances.

연구 동기 및 목표

  • 고도로 발전한 음성 변환 기술이 개인을 모방할 수 있는 개인정보 유출 문제를 해결하기 위해.
  • 화자 음성을 변환 공격에 저항하도록 만들기 위해, 비밀번호 없는 음성 변환을 방지하는 방어 메커니즘을 개발하기 위해.
  • 특히 분류 모델에 비해 연구가 부족한 생성 모델, 특히 음성 변환에 대한 적대적 공격를 탐구하기 위해.
  • 실제(블랙박스) 시나리오에서 객관적 화자 확인 및 주관적 听음 테스트를 통해 효과성을 평가하기 위해.

제안 방법

  • 음성 변환 모델이 처리하기 전에 화자의 발화에 인간이 인지하지 못하는 적대적 노이즈를 주입하기 위해.
  • 엔드 투 엔드 공격, 임bedding 공격, 피드백 공격의 세 가지 공격 전략을 제안하여 음성 변환 파이프라인의 서로 다른 구성 요소를 공격하기 위해.
  • 내용 및 화자 인코더를 별도로 가진 표준 인코더-디코더 아키텍처를 사용하고, 화자 인코더의 입력을 변형하여 화자 표현을 변경하기 위해.
  • 방어된 화자와 변환된 출력 간의 화자 임베딩 공간에서의 차이를 최대화하는 손실 함수를 사용하여 노이즈를 최적화하기 위해.
  • 실제 배포를 시뮬레이션하기 위해 화이트박스(모델 전체 액세스 가능) 및 블랙박스(프록시 모델 사용) 설정에서 공격를 적용하기 위해.
  • 객관적 지표(예: 화자 확인 정확도)와 인간 평가자들이 참여하는 주관적 听음 테스트를 통해 공격를 검증하기 위해.
Fig. 1 : The encoder-decoder based voice conversion model and the three proposed approaches. Perturbations are updated on the utterances providing speaker characteristics, as the blue dashed lines indicate.
Fig. 1 : The encoder-decoder based voice conversion model and the three proposed approaches. Perturbations are updated on the utterances providing speaker characteristics, as the blue dashed lines indicate.

실험 결과

연구 질문

  • RQ1인간이 인지하지 못하는 노이즈를 사용하여 제로샷 음성 변환 모델을 교란시킬 수 있으며, 원래 발화의 청각적 품질은 손상되지 않을 수 있는가?
  • RQ2객관적 및 주관적 평가 모두에서 변환된 음성이 방어된 화자와 다르게 들리는지에 대해 적대적 공격의 효과는 어떠한가?
  • RQ3공격자가 대상 모델에 대한 지식이 제한된 블랙박스 시나리오에서도 공격가 효과적인가?
  • RQ4인간 청취자들이 적대적 출력을 원래 화자와 다르게 인지하는 정도는 어떠한가?
  • RQ5Chou의 모델과 AutoVC와 같은 최신 음성 변환 모델에 대해 이 공격이 일관된 결과를 낼 수 있는가?

주요 결과

  • 화이트박스 시나리오에서, Chou의 모델에 대해 적대적 출력의 화자 확인 정확도는 1.5%로 떨어졌고, AutoVC에 대해서는 2.5%로 나타나 방어된 화자로부터의 심각한 이탈이 확인되었다.
  • 주관적 평가 결과, 58% 이상의 청취자가 적대적 출력을 다른 화자에서 나온 것으로 판단했으며(유형 I), 70% 이상이 적대적 입력이 원래 화자와 동일하다고 인지했다.
  • 블랙박스 공격도 효과적이었으며, Chou의 모델에 대해 12.0%, AutoVC에 대해 19.5%의 화자 확인 점수로 변환된 음성이 다른 화자에서 나온 것으로 나타났다.
  • AutoVC의 경우, 원래 출력이 같은 화자로 인지된 청취자는 27%에 불과하여 객관적 지표와 인간 인지 간의 괴리가 있음을 시사했다.
  • 적대적 입력은 음성 품질을 유지했으며, 객관적 및 주관적 테스트 모두에서 진짜 발화와 구분되지 않았다.
  • 이 방법은 다양한 모델과 설정에서 뛰어난 강건성을 보였으며, 화이트박스 및 블랙박스 평가 모두에서 일관된 성능을 보였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.