[Paper Review] Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
The paper introduces Emotion-Aware Prefix with Deep-Prefix Prompting to enable explicit emotion control in a two-stage VEVO-based voice conversion system, achieving substantial gains in Emotion Conversion Accuracy (ECA) while preserving speaker identity and quality.
Recent advances in zero-shot voice conversion have exhibited potential in emotion control, yet the performance is suboptimal or inconsistent due to their limited expressive capacity. We propose Emotion-Aware Prefix for explicit emotion control in a two-stage voice conversion backbone. We significantly improve emotion conversion performance, doubling the baseline Emotion Conversion Accuracy (ECA) from 42.40% to 85.50% while maintaining linguistic integrity and speech quality, without compromising speaker identity. Our ablation study suggests that a joint control of both sequence modulation and acoustic realization is essential to synthesize distinct emotions. Furthermore, comparative analysis verifies the generalizability of proposed method, while it provides insights on the role of acoustic decoupling in maintaining speaker identity.
Motivation & Objective
- Motivate explicit emotion control in zero-shot voice conversion to enhance expressiveness without sacrificing linguistic content or speaker identity.
- Extend VEVO with a content-invariant emotion prefix to steer sequence modulation.
- Investigate the hierarchical impact of emotion prompts across sequence modulation and acoustic realization stages.
- Assess the generalizability and the role of acoustic decoupling for identity preservation in emotion-controlled VC.
Proposed method
- Extend VEVO by adding an Emotion-Aware Prefix encoder that extracts an utterance-level emotion embedding from a reference mel-spectrogram.
- Use a Temporal-Shuffle Transformer, a Perceiver layer, and an Emotion Fusion Layer to produce a fixed-length emotion prefix E.
- Implement Deep-Prefix Prompting to inject E as layerwise KV-cache in the autoregressive token generator for sequence modulation.
- Condition the Acoustic Realization stage on reference audio tokens and ground-truth mel-spectrogram to realize final speech with preserved speaker identity.
- Fine-tune only the Emotion-Aware Prefix Encoder and apply LoRA to the AR Transformer for lightweight adaptation while keeping the backbone frozen.
- Train on the Emotion Speech Dataset (ESD) with 10 speakers across 5 emotions, using 300 training utterances per speaker-emotion pair.
Experimental results
Research questions
- RQ1Can explicit emotion control be achieved in a two-stage voice conversion framework by introducing an emotion-aware prefix?
- RQ2What is the relative contribution of sequence-level modulation versus acoustic realization in driving emotion conversion performance?
- RQ3Does acoustic decoupling help preserve speaker identity when adding explicit emotion control?
- RQ4How does the proposed method perform compared to VEVO and other baselines in objective and subjective measures of emotion, quality, and identity?
Key findings
- Emotion Conversion Accuracy (ECA) improves from 42.40% (VEVO) to 85.50% with the proposed method.
- Deep-Prefix Prompting further enhances ECA and emotion similarity (Emo SIM) without sacrificing quality or intelligibility.
- Sequence modulation is the primary driver of high-level emotion, with joint control across stages yielding the largest non-additive gains.
- Acoustic decoupling is beneficial for preserving speaker identity, as methods lacking a separate acoustic realization stage show stronger identity degradation.
- Subjective evaluations show improved emotion similarity and speaker preference for the proposed method (MOS and ABX tests).
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.