[Paper Review] The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods
The paper presents VCC 2018, introducing Hub (parallel) and Spoke (non-parallel) voice conversion tasks, a large crowdsourced perceptual evaluation, and analysis of both traditional and neural VC approaches, with N10 performing best on naturalness and similarity.
We present the Voice Conversion Challenge 2018, designed as a follow up to the 2016 edition with the aim of providing a common framework for evaluating and comparing different state-of-the-art voice conversion (VC) systems. The objective of the challenge was to perform speaker conversion (i.e. transform the vocal identity) of a source speaker to a target speaker while maintaining linguistic information. As an update to the previous challenge, we considered both parallel and non-parallel data to form the Hub and Spoke tasks, respectively. A total of 23 teams from around the world submitted their systems, 11 of them additionally participated in the optional Spoke task. A large-scale crowdsourced perceptual evaluation was then carried out to rate the submitted converted speech in terms of naturalness and similarity to the target speaker identity. In this paper, we present a brief summary of the state-of-the-art techniques for VC, followed by a detailed explanation of the challenge tasks and the results that were obtained.
Motivation & Objective
- Provide a common framework to evaluate and compare state-of-the-art voice conversion systems.
- Assess parallel and non-parallel VC methods under unified listening tests.
- Analyze the relationship between perceptual quality and intelligibility, and relate to ASV spoofing considerations.
Proposed method
- Describe the Hub task using parallel data with 4 source and 4 target speakers and 16 source–target pairs.
- Describe the Spoke task using non-parallel data with the same target speakers but different sources and utterances.
- Use a large crowdsourced listening test to rate naturalness and similarity of converted speech.
- Provide baseline systems (sprocket and Merlin) and document participant systems and vocoders used.
- Present an analysis of WER (ASR-based intelligibility) on converted speech to complement perceptual results.
Experimental results
Research questions
- RQ1How do parallel and non-parallel VC systems compare under the same evaluation framework?
- RQ2What are the perceptual naturalness and speaker similarity levels achievable with current VC approaches, including neural vocoders like WaveNet?
- RQ3What is the relationship between subjective quality (MOS) and objective intelligibility (WER) in VC outputs?
- RQ4Do VC submissions pose spoofing risks, and how do they relate to ASV countermeasures?
Key findings
- Twenty-three teams submitted Hub task systems, with 11 also participating in the Spoke task.
- N10 achieved the best naturalness, close to target speech, and high similarity across Hub and Spoke tasks.
- A WaveNet-based system (N10) delivered naturalness around 4.1 on a 5-point scale and about 80% of samples judged as the target speaker.
- Spoke (non-parallel) Task showed overall lower naturalness than Hub, reflecting greater task difficulty, while some systems still achieved reasonable similarity.
- There is a strong negative correlation between MOS (naturalness) and WER, indicating spectral distortions affect both perceptual quality and intelligibility.
- Baseline sprocket systems performed competitively in some same-gender cases but struggled in cross-gender conditions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.