[論文レビュー] Controllable cross-speaker emotion transfer for end-to-end speech synthesis.
本稿では、2つの感情分離モジュールを用いて感情と話者IDを分離することで、制御可能なクロス・スピーカー感情転送を可能にする、Tacotron2に基づくエンド・ツー・エンド音声合成フレームワークを提案する。感情強度を制御可能なスカラを導入し、話者漏れを低減させつつ、プロソディが多様で感情表現豊かな音声を、ターゲット話者に対して生成する。
The cross-speaker emotion transfer task in TTS particularly aims to synthesize speech for a target speaker with the emotion transferred from reference speech recorded by another (source) speaker. During the emotion transfer process, the identity information of the source speaker could also affect the synthesized results, resulting in the issue of speaker leakage. This paper proposes a new method with the aim to synthesize controllable emotional expressive speech and meanwhile maintain the target speaker's identity in the cross-speaker emotion TTS task. The proposed method is a Tacotron2-based framework with the emotion embedding as the conditioning variable to provide emotion information. Two emotion disentangling modules are contained in our method to 1) get speaker-independent and emotion-discriminative embedding, and 2) explicitly constrain the emotion and speaker identity of synthetic speech to be that as expected. Moreover, we present an intuitive method to control the emotional strength in the synthetic speech for the target speaker. Specifically, the learned emotion embedding is adjusted with a flexible scalar value, which allows controlling the emotion strength conveyed by the embedding. Extensive experiments have been conducted on a Mandarin disjoint corpus, and the results demonstrate that the proposed method is able to synthesize reasonable emotional speech for the target speaker. Compared to the state-of-the-art reference embedding learned methods, our method gets the best performance on the cross-speaker emotion transfer task, indicating that our method achieves the new state-of-the-art performance on learning the speaker-independent emotion embedding. Furthermore, the strength ranking test and pitch trajectories plots demonstrate that the proposed method can effectively control the emotion strength, leading to prosody-diverse synthetic speech.
研究の動機と目的
- エンド・ツー・エンドTTSにおける感情と話者IDの分離により、クロス・スピーカー感情転送における話者漏れを低減すること。
- ターゲット話者に対して合成音声の感情表現の強度を正確に制御できること。
- 話者に依存しない感情埋め込みを学習することで、感情音声合成におけるプロソディの質と多様性を向上させること。
- ターゲット話者IDを維持しつつ、クロス・スピーカー感情転送で最先端の性能を達成すること。
提案手法
- フレームワークはTacotron2に基づき、感情埋め込みを条件変数として用い、感情表現音声合成をガイドする。
- 2つの感情分離モジュールを導入:1つは話者に依存しない感情特徴を抽出し、もう1つは出力の感情と話者IDを明示的に制約する。
- 感情埋め込みに学習可能なスカラを適用し、合成音声における感情表現の強度を制御する。
- 生成音声の話者埋め込みがターゲット話者と一致するように明示的に制約することで、アイデンティティの一貫性を確保する。
- エンド・ツー・エンドで中国語の非重複コーパス上で訓練し、感情表現の豊かさと話者IDの保持の両方を最適化する。
- スカラ制御機構による感情強度の変調により、プロソディが多様な出力を実現する。
実験結果
リサーチクエスチョン
- RQ1ソース話者からターゲット話者への感情転送を、話者漏れを最小限に抑えて効果的に行うことができるか?
- RQ2感情表現の強度を、分離的かつ直感的な方法で制御できるか?
- RQ3話者に依存しない感情埋め込みを効果的に学習できるか?
- RQ4提案手法は、最先端のリファレンス埋め込みベースの手法よりも、クロス・スピーカー感情転送で優れた性能を発揮するか?
主な発見
- 提案手法は、既存のリファレンス埋め込みベース手法と比較して、クロス・スピーカー感情転送タスクで最先端の性能を達成した。
- 強度ランクテストにより、スカラ制御機構が合成発話全体における感情強度を効果的に調整していることが確認された。
- ピッチトレースのプロットから、本手法が多様な感情表現を持つプロソディを生成することが示された。
- 2つの感情分離モジュールにより、感情と話者IDを分離することで、話者漏れが著しく低減された。
- モデルはターゲット話者に対して妥当で自然な感情表現音声を合成し、高い話者アイデンティティ忠実度を維持した。
- 調整可能な感情強度を備えた、感情表現豊かな音声の生成において、本手法は頑健性と制御性を示した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。