Skip to main content
QUICK REVIEW

[Paper Review] The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods

Jaime Lorenzo-Trueba, Junichi Yamagishi|arXiv (Cornell University)|Apr 12, 2018
Speech Recognition and Synthesis3 references66 citations
TL;DR

The paper presents VCC 2018, introducing Hub (parallel) and Spoke (non-parallel) voice conversion tasks, a large crowdsourced perceptual evaluation, and analysis of both traditional and neural VC approaches, with N10 performing best on naturalness and similarity.

ABSTRACT

We present the Voice Conversion Challenge 2018, designed as a follow up to the 2016 edition with the aim of providing a common framework for evaluating and comparing different state-of-the-art voice conversion (VC) systems. The objective of the challenge was to perform speaker conversion (i.e. transform the vocal identity) of a source speaker to a target speaker while maintaining linguistic information. As an update to the previous challenge, we considered both parallel and non-parallel data to form the Hub and Spoke tasks, respectively. A total of 23 teams from around the world submitted their systems, 11 of them additionally participated in the optional Spoke task. A large-scale crowdsourced perceptual evaluation was then carried out to rate the submitted converted speech in terms of naturalness and similarity to the target speaker identity. In this paper, we present a brief summary of the state-of-the-art techniques for VC, followed by a detailed explanation of the challenge tasks and the results that were obtained.

Motivation & Objective

  • Provide a common framework to evaluate and compare state-of-the-art voice conversion systems.
  • Assess parallel and non-parallel VC methods under unified listening tests.
  • Analyze the relationship between perceptual quality and intelligibility, and relate to ASV spoofing considerations.

Proposed method

  • Describe the Hub task using parallel data with 4 source and 4 target speakers and 16 source–target pairs.
  • Describe the Spoke task using non-parallel data with the same target speakers but different sources and utterances.
  • Use a large crowdsourced listening test to rate naturalness and similarity of converted speech.
  • Provide baseline systems (sprocket and Merlin) and document participant systems and vocoders used.
  • Present an analysis of WER (ASR-based intelligibility) on converted speech to complement perceptual results.

Experimental results

Research questions

  • RQ1How do parallel and non-parallel VC systems compare under the same evaluation framework?
  • RQ2What are the perceptual naturalness and speaker similarity levels achievable with current VC approaches, including neural vocoders like WaveNet?
  • RQ3What is the relationship between subjective quality (MOS) and objective intelligibility (WER) in VC outputs?
  • RQ4Do VC submissions pose spoofing risks, and how do they relate to ASV countermeasures?

Key findings

  • Twenty-three teams submitted Hub task systems, with 11 also participating in the Spoke task.
  • N10 achieved the best naturalness, close to target speech, and high similarity across Hub and Spoke tasks.
  • A WaveNet-based system (N10) delivered naturalness around 4.1 on a 5-point scale and about 80% of samples judged as the target speaker.
  • Spoke (non-parallel) Task showed overall lower naturalness than Hub, reflecting greater task difficulty, while some systems still achieved reasonable similarity.
  • There is a strong negative correlation between MOS (naturalness) and WER, indicating spectral distortions affect both perceptual quality and intelligibility.
  • Baseline sprocket systems performed competitively in some same-gender cases but struggled in cross-gender conditions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.