Skip to main content
QUICK REVIEW

[Paper Review] StarGAN-VC2: Rethinking Conditional Methods for StarGAN-Based Voice Conversion

Takuhiro Kaneko, Hirokazu Kameoka|arXiv (Cornell University)|Jul 29, 2019
Speech Recognition and Synthesis55 references17 citations
TL;DR

StarGAN-VC2 proposes an improved voice conversion framework that rethinks conditional methods in StarGAN-based models by introducing a source-and-target conditional adversarial loss and a modulation-based network architecture. These enhancements significantly improve speech quality and speaker similarity, outperforming StarGAN-VC in both objective and subjective evaluations on multi-speaker voice conversion tasks.

ABSTRACT

Non-parallel multi-domain voice conversion (VC) is a technique for learning mappings among multiple domains without relying on parallel data. This is important but challenging owing to the requirement of learning multiple mappings and the non-availability of explicit supervision. Recently, StarGAN-VC has garnered attention owing to its ability to solve this problem only using a single generator. However, there is still a gap between real and converted speech. To bridge this gap, we rethink conditional methods of StarGAN-VC, which are key components for achieving non-parallel multi-domain VC in a single model, and propose an improved variant called StarGAN-VC2. Particularly, we rethink conditional methods in two aspects: training objectives and network architectures. For the former, we propose a source-and-target conditional adversarial loss that allows all source domain data to be convertible to the target domain data. For the latter, we introduce a modulation-based conditional method that can transform the modulation of the acoustic feature in a domain-specific manner. We evaluated our methods on non-parallel multi-speaker VC. An objective evaluation demonstrates that our proposed methods improve speech quality in terms of both global and local structure measures. Furthermore, a subjective evaluation shows that StarGAN-VC2 outperforms StarGAN-VC in terms of naturalness and speaker similarity. The converted speech samples are provided at http://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/stargan-vc2/index.html.

Motivation & Objective

  • To address the performance gap between real and converted speech in non-parallel multi-domain voice conversion.
  • To improve the generalization and fidelity of StarGAN-VC by rethinking conditional methods in training objectives and network design.
  • To enable high-quality, single-generator voice conversion across multiple speakers without parallel training data.
  • To enhance both global (e.g., spectral structure) and local (e.g., modulation dynamics) feature alignment in converted speech.
  • To achieve better naturalness and speaker similarity in subjective listening tests compared to prior state-of-the-art models.

Proposed method

  • Proposes a source-and-target conditional adversarial loss that encourages all source-domain speech to be converted toward the target-domain distribution, improving discriminative generalization.
  • Introduces a modulation-based conditional mechanism that applies domain-specific modulation to acoustic features, enabling fine-grained control over spectral dynamics.
  • Uses a single generator with domain-conditional conditioning, maintaining the scalability of StarGAN-VC while improving feature fidelity.
  • Employs a GAN-based framework with identity loss and cycle consistency, adapted to multi-speaker settings using domain codes.
  • Applies Mel-cepstral distortion (MCD) and modulation spectra distance (MSD) as objective metrics to evaluate global and local feature similarity.
  • Conducts XAB preference tests and MOS evaluations to assess naturalness and speaker similarity in subjective listening tests.

Experimental results

Research questions

  • RQ1Can a restructured adversarial loss improve the alignment of converted speech with target speech in multi-speaker voice conversion?
  • RQ2Does a modulation-based conditional network architecture enhance local spectral structure preservation in non-parallel voice conversion?
  • RQ3How do the proposed conditional methods compare to conventional channel-wise conditioning and standard adversarial losses in terms of speech quality and speaker similarity?
  • RQ4To what extent does the proposed method reduce the gap between real and converted speech in both objective and subjective evaluations?
  • RQ5Can the improved model generalize beyond multi-speaker conversion to other multi-domain voice conversion tasks?

Key findings

  • StarGAN-VC2 achieved a Mel-cepstral distortion (MCD) of 6.90 dB and modulation spectra distance (MSD) of 1.89 dB, outperforming StarGAN-VC in both global and local feature similarity.
  • The proposed source-and-target conditional adversarial loss reduced MCD by 0.21 dB and MSD by 0.52 dB compared to the baseline StarGAN-VC with standard loss.
  • The modulation-based conditional network reduced MSD by 0.66 dB compared to the channel-wise method, indicating superior local structure preservation.
  • Subjective MOS for naturalness was significantly higher for StarGAN-VC2 (p < 0.05), with a mean score of 3.85 compared to 3.67 for StarGAN-VC.
  • Preference scores for speaker similarity were 72% for StarGAN-VC2 versus 60% for StarGAN-VC, indicating a statistically significant improvement in speaker identity preservation.
  • The model demonstrated consistent superiority across all categories: full conversion, intra-gender, and inter-gender conversion.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.