[论文解读] Controllable cross-speaker emotion transfer for end-to-end speech synthesis.
该论文提出了一种基于Tacotron2的端到端语音合成框架,通过两个情感解耦模块解耦情感与说话人身份,实现了可控的跨说话人情感迁移。该方法引入可学习标量以控制情感强度,实现了最先进的性能,同时减少了说话人泄漏,并生成了针对目标说话人的多样语调、情感丰富的语音。
The cross-speaker emotion transfer task in TTS particularly aims to synthesize speech for a target speaker with the emotion transferred from reference speech recorded by another (source) speaker. During the emotion transfer process, the identity information of the source speaker could also affect the synthesized results, resulting in the issue of speaker leakage. This paper proposes a new method with the aim to synthesize controllable emotional expressive speech and meanwhile maintain the target speaker's identity in the cross-speaker emotion TTS task. The proposed method is a Tacotron2-based framework with the emotion embedding as the conditioning variable to provide emotion information. Two emotion disentangling modules are contained in our method to 1) get speaker-independent and emotion-discriminative embedding, and 2) explicitly constrain the emotion and speaker identity of synthetic speech to be that as expected. Moreover, we present an intuitive method to control the emotional strength in the synthetic speech for the target speaker. Specifically, the learned emotion embedding is adjusted with a flexible scalar value, which allows controlling the emotion strength conveyed by the embedding. Extensive experiments have been conducted on a Mandarin disjoint corpus, and the results demonstrate that the proposed method is able to synthesize reasonable emotional speech for the target speaker. Compared to the state-of-the-art reference embedding learned methods, our method gets the best performance on the cross-speaker emotion transfer task, indicating that our method achieves the new state-of-the-art performance on learning the speaker-independent emotion embedding. Furthermore, the strength ranking test and pitch trajectories plots demonstrate that the proposed method can effectively control the emotion strength, leading to prosody-diverse synthetic speech.
研究动机与目标
- 通过在端到端TTS中解耦情感与说话人身份,解决跨说话人情感迁移中的说话人泄漏问题。
- 实现对目标说话人合成语音中情感表达强度的精确控制。
- 通过学习与说话人无关的情感嵌入,提升情感语音合成的语调质量和多样性。
- 在保持目标说话人身份的同时,实现跨说话人情感迁移的最先进性能。
提出的方法
- 该框架基于Tacotron2,使用情感嵌入作为条件变量,以引导情感语音合成。
- 引入两个情感解耦模块:一个用于提取与说话人无关、具有情感判别能力的嵌入,另一个用于显式约束输出的情感与说话人身份。
- 在情感嵌入上应用可学习标量,以控制合成语音中情感表达的强度。
- 通过显式约束生成语音的说话人嵌入以匹配目标说话人,确保身份一致性。
- 在普通话分离语料上端到端训练该架构,以优化情感表现力和说话人身份保持。
- 通过标量控制机制调节情感强度,实现语调多样的输出。
实验结果
研究问题
- RQ1TTS系统能否在最小化说话人泄漏的前提下,有效将情感从源说话人迁移到目标说话人?
- RQ2如何以解耦且直观的方式控制合成语音中的情感强度?
- RQ3能否有效学习与说话人无关的情感嵌入,以提升跨说话人情感迁移性能?
- RQ4所提方法在跨说话人情感迁移中是否优于现有的最先进参考嵌入方法?
主要发现
- 与现有基于参考嵌入的方法相比,所提方法在跨说话人情感迁移任务中实现了最先进性能。
- 强度排序测试证实,标量控制机制能有效调节合成语句中的情感强度。
- 基频轨迹图显示,该方法能生成语调多样的语音,具有不同的情感表达。
- 两个情感解耦模块成功通过将情感与说话人身份分离,减少了说话人泄漏。
- 该模型能为目标说话人生成合理且自然的情感语音,同时保持高度的说话人身份保真度。
- 该方法在生成情感丰富且可调节情感强度的语音方面表现出鲁棒性和可控性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。