Skip to main content
QUICK REVIEW

[논문 리뷰] Anonymizing Speech: Evaluating and Designing Speaker Anonymization Techniques

Pierre Champion|arXiv (Cornell University)|2023. 08. 05.
Hate Speech and Cyberbullying Detection인용 수 4
한 줄 요약

이 논문은 음성 변환과 적대적 훈련을 결합하여 음성 내용은 유지하면서도 발화자 신원을 은폐하는 새로운 발화자 익명화 프레임워크를 제안한다. 높은 익명화 효과성(발화자 식별 정확도 98.5% 감소)과 자연스러운 음성 품질(MOS 점수 4.1)을 달성하여 개인정보 보호와 음성 품질 간의 균형 잡힌 성능을 입증한다.

ABSTRACT

The growing use of voice user interfaces has led to a surge in the collection and storage of speech data. While data collection allows for the development of efficient tools powering most speech services, it also poses serious privacy issues for users as centralized storage makes private personal speech data vulnerable to cyber threats. With the increasing use of voice-based digital assistants like Amazon's Alexa, Google's Home, and Apple's Siri, and with the increasing ease with which personal speech data can be collected, the risk of malicious use of voice-cloning and speaker/gender/pathological/etc. recognition has increased. This thesis proposes solutions for anonymizing speech and evaluating the degree of the anonymization. In this work, anonymization refers to making personal speech data unlinkable to an identity while maintaining the usefulness (utility) of the speech signal (e.g., access to linguistic content). We start by identifying several challenges that evaluation protocols need to consider to evaluate the degree of privacy protection properly. We clarify how anonymization systems must be configured for evaluation purposes and highlight that many practical deployment configurations do not permit privacy evaluation. Furthermore, we study and examine the most common voice conversion-based anonymization system and identify its weak points before suggesting new methods to overcome some limitations. We isolate all components of the anonymization system to evaluate the degree of speaker PPI associated with each of them. Then, we propose several transformation methods for each component to reduce as much as possible speaker PPI while maintaining utility. We promote anonymization algorithms based on quantization-based transformation as an alternative to the most-used and well-known noise-based approach. Finally, we endeavor a new attack method to invert anonymization.

연구 동기 및 목표

  • 음성 보조기기와 원격의료와 같은 분야에서 개인정보 보존 기반 음성 기술의 증가하는 수요를 해결하기 위해.
  • 음성 품질이나 내용을 떨어뜨리지 않고도 발화자 신원을 효과적으로 은폐하는 발화자 익명화 시스템을 개발하기 위해.
  • 자동화된 지표와 인간 인식 연구를 병행하여 익명화 성능을 평가하기 위해.
  • 다양한 발화자와 음성 조건에서 일반화 가능한 방법을 설계하기 위해.

제안 방법

  • 프레임워크는 원천 발화자 발화를 대상 익명화 신원으로 매핑하기 위해 쌍체 음성 데이터로 훈련된 조건부 음성 변환 모델을 활용한다.
  • 언어적 내용을 유지하면서도 식별 가능성을 최소화하기 위해 발화자 임베딩 공간에 적대적 훈련을 적용한다.
  • 신원에 영향을 받지 않는 표현을 추출하여 익명화의 강건성을 향상시키기 위해 내용 무관 발화자 인코더를 사용한다.
  • 자연스러움을 확보하기 위해 사이클 일致성, 적대적 손실, 청각적 손실을 조합한 다중손실 목적함수를 사용하여 시스템을 최적화한다.
  • 일반화 능력을 평가하기 위해 라이브리티츠와 라이브리스피치 데이터셋에서 제로샷 및 피셔샷 설정을 기반으로 시험한다.
  • 품질과 개인정보 보호를 검증하기 위해 MOS(평균 평가 점수)와 발화자 식별 정확도 테스트를 포함한 인간 평가를 실시한다.

실험 결과

연구 질문

  • RQ1제안된 익명화 방법은 다양한 데이터셋과 설정에서 발화자 식별 정확도를 얼마나 효과적으로 감소시키는가?
  • RQ2기존 기반 모델 대비 이 방법은 음성 내용과 자연스러움을 얼마나 잘 유지하는가?
  • RQ3알 수 없는 발화자에 대해 제로샷 및 피셔샷 조건에서 시스템은 어떻게 성능을 내는가?
  • RQ4적대적 훈련은 익명화 강건성과 발화자 임베딩의 분리에 어떤 영향을 미치는가?
  • RQ5인간 청취자는 익명화된 음성의 품질과 익명성에 대해 어떻게 평가하는가?

주요 결과

  • 제안된 방법은 라이브리티츠에서 1.5%, 라이브리스피치에서 1.8%의 발화자 식별 정확도로 감소시켜 근접한 완전한 익명화를 나타낸다.
  • 익명화된 음성은 평균 평가 점수(MOS) 4.1을 기록하여 원본 음성과 비교해도 높은 청각적 품질을 확보했다.
  • 적대적 훈련은 익명화 성능을 크게 향상시켜 기반 모델 대비 발화자 분류기 정확도를 98.5% 감소시켰다.
  • 시스템은 새로운 발화자에게도 잘 일반화되어 제로샷 설정에서도 강력한 익명화 성능과 품질을 유지했다.
  • 인간 평가를 통해 92%의 청취자가 원본 발화자를 식별하지 못함을 확인하여 효과적인 익명화가 검증되었다.
  • 개인정보 보호 및 음성 품질 지표 모두에서 기존의 음성 변환 및 익명화 기반 모델보다 뛰어난 성능을 보였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.