[論文レビュー] MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis
MsEmoTTSは、感情をグローバル(カテゴリ)、発話(プロソディーパatters)、局所的(音節レベルの強度)の3スケールでモデル化するマルチスケールな感情音声合成フレームワークを提案する。3つの専用モジュールを用いて、リファレンス音声の転送、テキストベースの感情予測、手動による感情強度制御を可能とし、中国語コーパスにおける感情転送および予測タスクで先行手法を上回る性能を達成した。
Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in recent years. Previous methods performed the expressive speech synthesis either with explicit labels or with a fixed-length style embedding extracted from reference audio, both of which can only learn an average style and thus ignores the multi-scale nature of speech prosody. In this paper, we propose MsEmoTTS, a multi-scale emotional speech synthesis framework, to model the emotion from different levels. Specifically, the proposed method is a typical attention-based sequence-to-sequence model and with proposed three modules, including global-level emotion presenting module (GM), utterance-level emotion presenting module (UM), and local-level emotion presenting module (LM), to model the global emotion category, utterance-level emotion variation, and syllable-level emotion strength, respectively. In addition to modeling the emotion from different levels, the proposed method also allows us to synthesize emotional speech in different ways, i.e., transferring the emotion from reference audio, predicting the emotion from input text, and controlling the emotion strength manually. Extensive experiments conducted on a Chinese emotional speech corpus demonstrate that the proposed method outperforms the compared reference audio-based and text-based emotional speech synthesis methods on the emotion transfer speech synthesis and text-based emotion prediction speech synthesis respectively. Besides, the experiments also show that the proposed method can control the emotion expressions flexibly. Detailed analysis shows the effectiveness of each module and the good design of the proposed method.
研究の動機と目的
- 既存の表現力のあるTTSシステムにおける単一スケールの感情モデリングの限界を解消し、人間のプロソディの多様なスケール的性質を捉えること。
- リファレンスベースの感情転送、テキストベースの感情予測、手動による感情強度制御の3つの異なるモードを通じて、柔軟な感情付き音声合成を実現すること。
- グローバルな感情カテゴリ、発話レベルのプロソディック変動、音節レベルの感情強度という3段階の階層的レベルで感情をモデリングすること。
- 異なる音声ユニットにおける微細な感情の変動を捉えることで、合成音声の表現力と多様性を向上させること。
提案手法
- 本モデルは、3つの専用モジュール(グローバルレベル感情表現モジュール: GM、発話レベル感情表現モジュール: UM、局所レベル感情表現モジュール: LM)を備えたアテンションベースのシーケンス・ツー・シーケンスアーキテクチャを採用する。
- GMは事前学習済みの感情分類器を用いて発話全体の感情カテゴリを分類し、グローバルなスタイル信号を提供する。
- UMは発話内でのプロソディックパターン(例:イントネーションカーブ)を、発話全体にわたるフォルクレーターとエネルギーの変動をモデル化することで学習する。
- LMは音節レベルで感情強度値を割り当て、局所的な感情強度を制御する。実験では中国語において音節がフォノームよりも適していることが示された。
- リファレンス音声またはテキストからの感情表現は、アダプティブレイヤーナルムライゼーションを介してTTSデコーダーに統合され、スタイル転送と制御が可能になる。
- マルチタスク学習を用いたエンドツーエンド学習が可能であり、感情分類と強度予測が同時に最適化される。
実験結果
リサーチクエスチョン
- RQ1グローバル、発話、局所の3スケールで感情をモデリングすることで、合成された感情付き音声の表現力と自然さが向上するか?
- RQ2本手法が提示するマルチスケール感情表現は、単一スケールのリファレンスベースまたはラベルベースの手法と比較して、感情転送および予測タスクで優れているか?
- RQ3中国語の音声において、音節レベルの感情強度を用いることで、フォノームレベルや文レベルの表現と比較して合成品質が向上するか?
- RQ4感情カテゴリと強度を手動で入力することで、どの程度の柔軟な感情制御が可能になるか?
主な発見
- MsEmoTTSは中国語の感情付き音声コーパスにおいて、リファレンスベースおよびテキストベースの感情付きTTS手法を、感情転送および感情予測タスクの両方で上回った。
- 音節レベルの感情強度を用いた場合、Mel-cepstral distortion (MCD) は4.11 dBに低下し、文レベルの強度を用いた場合の4.65 dBと比較して顕著に低いことが確認された。これは微細な局所モデリングの有効性を裏付けている。
- 音節レベルの感情強度手法は、フォノームレベルの手法と比較してMCD低減率が1.1%向上し、トーン言語としての中国語に適していることが示された。
- メルスペクトログ램およびf0カーブの可視化分析により、発話レベルモジュール(UM)がリファレンス音声からのフォルクレーター変動トレンドを効果的に捉え、転送していることが確認された。
- グローバルな感情カテゴリと局所的な感情強度を手動で制御することで、望ましい感情表現の正確な合成が可能であり、モデルの柔軟性が示された。
- GMとUMに異なる感情カテゴリのリファレンスが同時に使用された場合(例:矛盾する感情カテゴリ)、合成音声に不自然なハイブリッド感情が現れることが観察され、グローバルおよび発話レベルのプロソディの間の依存関係が示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。