Skip to main content
QUICK REVIEW

[Paper Review] Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2

Suvendu Sekhar Mohanty|arXiv (Cornell University)|Mar 12, 2026
Emotion and Mood Recognition0 citations
TL;DR

The paper introduces a causal prosody mediation framework that augments FastSpeech2 with emotion conditioning and counterfactual training to disentangle emotion from linguistic content, improving expressive TTS prosody and controllability.

ABSTRACT

We propose a novel causal prosody mediation framework for expressive text-to-speech (TTS) synthesis. Our approach augments the FastSpeech2 architecture with explicit emotion conditioning and introduces counterfactual training objectives to disentangle emotional prosody from linguistic content. By formulating a structural causal model of how text (content), emotion, and speaker jointly influence prosody (duration, pitch, energy) and ultimately the speech waveform, we derive two complementary loss terms: an Indirect Path Constraint (IPC) to enforce that emotion affects speech only through prosody, and a Counterfactual Prosody Constraint (CPC) to encourage distinct prosody patterns for different emotions. The resulting model is trained on multi-speaker emotional corpora (LibriTTS, EmoV-DB, VCTK) with a combined objective that includes standard spectrogram reconstruction and variance prediction losses alongside our causal losses. In evaluations on expressive speech synthesis, our method achieves significantly improved prosody manipulation and emotion rendering, with higher mean opinion scores (MOS) and emotion accuracy than baseline FastSpeech2 variants. We also observe better intelligibility (low WER) and speaker consistency when transferring emotions across speakers. Extensive ablations confirm that the causal objectives successfully separate prosody attribution, yielding an interpretable model that allows controlled counterfactual prosody editing (e.g. "same utterance, different emotion") without compromising naturalness. We discuss the implications for identifiability in prosody modeling and outline limitations such as the assumption that emotion effects are fully captured by pitch, duration, and energy. Our work demonstrates how integrating causal learning principles into TTS can improve controllability and expressiveness in generated speech.

Motivation & Objective

  • Motivate expressive TTS by disentangling emotion from linguistic content in prosody.
  • Develop a structural causal model linking text, emotion, speaker, and prosody to speech.
  • Introduce loss terms that enforce causal constraints and enable counterfactual prosody editing.

Proposed method

  • Augment FastSpeech2 with explicit emotion conditioning.
  • Formulate a structural causal model where text, emotion, and speaker influence prosody (duration, pitch, energy) and waveform.
  • Derive two loss terms: Indirect Path Constraint (IPC) and Counterfactual Prosody Constraint (CPC).
  • Train on multi-speaker emotional corpora (LibriTTS, EmoV-DB, VCTK) with standard spectrogram and variance losses plus causal losses.
  • Evaluate prosody manipulation, emotion rendering, intelligibility (WER), and speaker consistency via ablations.

Experimental results

Research questions

  • RQ1Can emotion effects on prosody be fully captured by duration, pitch, and energy under a causal framework?
  • RQ2Do IPC and CPC constraints enable emotion-specific prosody while preserving naturalness and intelligibility?
  • RQ3Does counterfactual editing (same utterance, different emotion) achieve perceptually distinct prosody without performance loss?
  • RQ4How does the model perform in cross-speaker emotion transfer in terms of MOS, emotion accuracy, and WER?

Key findings

  • The proposed causal objectives improve prosody manipulation and emotion rendering versus baseline FastSpeech2 variants.
  • The model achieves higher MOS and emotion accuracy than baselines.
  • Intelligibility remains high (low WER) and speaker consistency improves when transferring emotions across speakers.
  • Ablations show the causal losses successfully separate prosody attribution and enable interpretable counterfactual editing.
  • The approach discusses identifiability considerations and limitations, such as emotion effects being captured by pitch, duration, and energy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.