Skip to main content
QUICK REVIEW

[논문 리뷰] Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2

Suvendu Sekhar Mohanty|arXiv (Cornell University)|2026. 03. 12.
Emotion and Mood Recognition인용 수 0
한 줄 요약

이 논문은 FastSpeech2에 감정 조건화와 반사실적 학습을 더한 인과적 운율 매개 프레임워크를 도입하여 감정과 언어 내용의 구분을 개선하고 표현력 있는 TTS 운율 및 제어 가능성을 향상시킨다.

ABSTRACT

We propose a novel causal prosody mediation framework for expressive text-to-speech (TTS) synthesis. Our approach augments the FastSpeech2 architecture with explicit emotion conditioning and introduces counterfactual training objectives to disentangle emotional prosody from linguistic content. By formulating a structural causal model of how text (content), emotion, and speaker jointly influence prosody (duration, pitch, energy) and ultimately the speech waveform, we derive two complementary loss terms: an Indirect Path Constraint (IPC) to enforce that emotion affects speech only through prosody, and a Counterfactual Prosody Constraint (CPC) to encourage distinct prosody patterns for different emotions. The resulting model is trained on multi-speaker emotional corpora (LibriTTS, EmoV-DB, VCTK) with a combined objective that includes standard spectrogram reconstruction and variance prediction losses alongside our causal losses. In evaluations on expressive speech synthesis, our method achieves significantly improved prosody manipulation and emotion rendering, with higher mean opinion scores (MOS) and emotion accuracy than baseline FastSpeech2 variants. We also observe better intelligibility (low WER) and speaker consistency when transferring emotions across speakers. Extensive ablations confirm that the causal objectives successfully separate prosody attribution, yielding an interpretable model that allows controlled counterfactual prosody editing (e.g. "same utterance, different emotion") without compromising naturalness. We discuss the implications for identifiability in prosody modeling and outline limitations such as the assumption that emotion effects are fully captured by pitch, duration, and energy. Our work demonstrates how integrating causal learning principles into TTS can improve controllability and expressiveness in generated speech.

연구 동기 및 목표

  • 프로소디에서 감정을 언어적 내용과 구분하여 표현력 있는 TTS를 고무한다.
  • 텍스트, 감정, 화자, 운율을 음성으로 연결하는 구조적 인과 모델을 개발한다.
  • 인과 제약을 강제하고 반사실적 운율 편집을 가능하게 하는 손실 항을 도입한다.

제안 방법

  • 명시적 감정 조건화를 통해 FastSpeech2를 보강한다.
  • 텍스트, 감정, 화자가 운율(지속 시간, 피치, 에너지) 및 파형에 영향을 미치는 구조적 인과 모델을 구성한다.
  • 간접 경로 제약(IPC)과 반사실적 운율 제약(CPC) 두 손실 항을 도출한다.
  • 표준 스펙트로그램 및 분산 손실과 함께 인과 손실을 포함하여 다화자 감정 코퍼스(LibriTTS, EmoV-DB, VCTK)에서 학습한다.
  • 절제 실험(ablation)을 통해 운율 조작, 감정 렌더링, 가독성(WER), 화자 일관성을 평가한다.

실험 결과

연구 질문

  • RQ1인과 프레임워크 하에서 운율에 미치는 감정 효과가 지속 시간, 피치, 에너지로 완전히 포착될 수 있는가?
  • RQ2IPC와 CPC 제약이 자연스러움과 가독성을 보존하면서 감정 특이적 운율을 가능하게 하는가?
  • RQ3반사실적 편집(같은 음성 단편, 다른 감정)이 성능 저하 없이 지각적으로 구별되는 운율을 달성하는가?
  • RQ4MOS, 감정 정확도, WER 면에서 교차 화자 감정 전이에 대해 모델의 성능은 어떠한가?

주요 결과

  • 제안된 인과 목표는 기본 FastSpeech2 변형 대비 운율 조작 및 감정 렌더링을 향상시킨다.
  • 모델은 베이스라인보다 더 높은 MOS와 감정 정확도를 달성한다.
  • 가독성은 여전히 높고(WER 낮음) 화자 간 감정 전이 시 화자 일관성이 향상된다.
  • 절제 실험으로 인과 손실이 운율 속성 할당을 성공적으로 분리하고 해석 가능한 반사실 편집을 가능하게 함을 보인다.
  • 해당 방법은 식별성 고려사항과 한계를 논의하며, 예를 들어 감정 효과가 피치, 지속 시간, 에너지로 포착될 수 있음을 지적한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.