Skip to main content
QUICK REVIEW

[Paper Review] A Spoofing Benchmark for the 2018 Voice Conversion Challenge: Leveraging from Spoofing Countermeasures for Speech Artifact Assessment

Tomi Kinnunen, Jaime Lorenzo-Trueba|arXiv (Cornell University)|Apr 23, 2018
Speech Recognition and Synthesis6 references3 citations
TL;DR

This paper proposes using spoofing countermeasures (CMs) as an objective, reference-free method to assess speech artifacts in voice conversion (VC) systems. By training a constant-Q cepstral coefficient (CQCC)-based CM to distinguish real from converted speech, the equal error rate (EER) serves as a proxy for artifact severity—lower EER indicates more detectable artifacts. The key finding is that no VCC'18 system achieves a 50% EER (ideal), indicating all systems introduce measurable artifacts, with Wavenet-based systems showing high detectability despite strong perceptual quality.

ABSTRACT

Voice conversion (VC) aims at conversion of speaker characteristic without altering content. Due to training data limitations and modeling imperfections, it is difficult to achieve believable speaker mimicry without introducing processing artifacts; performance assessment of VC, therefore, usually involves both speaker similarity and quality evaluation by a human panel. As a time-consuming, expensive, and non-reproducible process, it hinders rapid prototyping of new VC technology. We address artifact assessment using an alternative, objective approach leveraging from prior work on spoofing countermeasures (CMs) for automatic speaker verification. Therein, CMs are used for rejecting `fake' inputs such as replayed, synthetic or converted speech but their potential for automatic speech artifact assessment remains unknown. This study serves to fill that gap. As a supplement to subjective results for the 2018 Voice Conversion Challenge (VCC'18) data, we configure a standard constant-Q cepstral coefficient CM to quantify the extent of processing artifacts. Equal error rate (EER) of the CM, a confusability index of VC samples with real human speech, serves as our artifact measure. Two clusters of VCC'18 entries are identified: low-quality ones with detectable artifacts (low EERs), and higher quality ones with less artifacts. None of the VCC'18 systems, however, is perfect: all EERs are < 30 % (the `ideal' value would be 50 %). Our preliminary findings suggest potential of CMs outside of their original application, as a supplemental optimization and benchmarking tool to enhance VC technology.

Motivation & Objective

  • To address the lack of standardized, objective metrics for evaluating voice conversion (VC) quality beyond subjective listening tests.
  • To explore whether spoofing countermeasures (CMs), originally designed for anti-spoofing in automatic speaker verification, can be repurposed for objective artifact assessment in VC.
  • To provide a reference-free, text-independent alternative to perceptual MOS scores for benchmarking VC systems.
  • To evaluate the correlation between CM-based EER and subjective quality ratings (MOS) in the VCC'18 challenge.
  • To identify which VC methods produce more or less detectable artifacts using CMs, especially for newer techniques like Wavenet.

Proposed method

  • A standard constant-Q cepstral coefficient (CQCC) front-end is used to extract acoustic features from converted speech and real human speech.
  • A Gaussian Mixture Model (GMM) backend is trained to classify input utterances as either bona fide (real) or spoofed (converted) speech.
  • The equal error rate (EER) of the trained CM is computed as a performance metric, where lower EER indicates higher detectability of artifacts.
  • The CM is evaluated on VCC'18 challenge entries using both the baseline and challenge system outputs, with EER serving as an objective artifact score.
  • The method is applied without access to source speaker waveforms or text transcripts, making it reference-free and text-independent.
  • The EER is compared with subjective MOS scores from the VCC'18 perceptual tests to assess complementarity between machine and human perception.

Experimental results

Research questions

  • RQ1Can spoofing countermeasures be effectively repurposed as an objective metric for assessing speech artifacts in voice conversion systems?
  • RQ2How does the EER of a CM trained on VCC'18 data correlate with subjective quality scores (MOS) from human listeners?
  • RQ3Which VC systems produce the most and least detectable artifacts according to the CM-based EER metric?
  • RQ4To what extent do modern VC techniques like Wavenet-based synthesis introduce artifacts that are detectable by CMs, even if they are perceptually natural?
  • RQ5Can the CM-based EER serve as a reliable, reproducible, and efficient alternative to time-consuming subjective evaluations in VC benchmarking?

Key findings

  • All VCC'18 systems produced detectable artifacts, as evidenced by EER values below the ideal 50%, with none achieving perfect fooling of the CM.
  • Systems using waveform filtering, SuperVP, and Griffin-Lim-based generation showed lower EERs (better performance), indicating fewer detectable artifacts.
  • Wavenet-based systems, despite high MOS scores, exhibited relatively high EERs, suggesting they introduce artifacts that are not perceptually obvious but are detectable by machine models.
  • The EER metric showed sensitivity to feature selection and model configuration, indicating that CM performance depends on training and hyperparameter choices.
  • There is no strong correlation between CM-based EER and subjective MOS, confirming that human and machine perception of speech quality differ significantly.
  • The study reveals a gap: current state-of-the-art VC systems can fool human listeners but not necessarily spoofing countermeasures, indicating room for improvement in artifact minimization.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.