Skip to main content
QUICK REVIEW

[論文レビュー] Effective Use of Variational Embedding Capacity in Expressive End-to-End Speech Synthesis

Eric Battenberg, Soroosh Mariooryad|arXiv (Cornell University)|Jun 8, 2019
Speech Recognition and Synthesis参考文献 24被引用数 43
ひとこと要約

Capacitronは埋め込み容量を変分後方をテキストと話者で条件付けることで多用途な転送と高品質な prior サンプリングを可能にする統一フレームワークとして提示し、階層的潜在変数が忠実度とばらつきを制御する。

ABSTRACT

Recent work has explored sequence-to-sequence latent variable models for expressive speech synthesis (supporting control and transfer of prosody and style), but has not presented a coherent framework for understanding the trade-offs between the competing methods. In this paper, we propose embedding capacity (the amount of information the embedding contains about the data) as a unified method of analyzing the behavior of latent variable models of speech, comparing existing heuristic (non-variational) methods to variational methods that are able to explicitly constrain capacity using an upper bound on representational mutual information. In our proposed model (Capacitron), we show that by adding conditional dependencies to the variational posterior such that it matches the form of the true posterior, the same model can be used for high-precision prosody transfer, text-agnostic style transfer, and generation of natural-sounding prior samples. For multi-speaker models, Capacitron is able to preserve target speaker identity during inter-speaker prosody transfer and when drawing samples from the latent prior. Lastly, we introduce a method for decomposing embedding capacity hierarchically across two sets of latents, allowing a portion of the latent variability to be specified and the remaining variability sampled from a learned prior. Audio examples are available on the web.

研究の動機と目的

  • 埋め込み容量(表現的相互情報に相当する量)を潜在変数TTSモデルを分析する統一的レンズとして Motivateする。
  • KL項が埋め込み容量を上界付けし、ラグランジュ乗数を用いて所望の容量をターゲットに制御できることを示す。
  • variational posteriorをテキストと話者情報で conditioning することで、マルチタスク転送(韻律転送、スタイル転送)と自然なpriorサンプルを実現する。
  • 高低レベルの変動を分離し、制御された転送とサンプリングを可能にする階層的潜在構造を導入する。

提案手法

  • 埋め込み容量をKL(q(z|x)||p(z))に制限される相互情報様の量として定義する。
  • 容量制約Cの下で平均埋め込み容量(R^avg)を境界づけるLagrangian目的関数 min_theta max_beta>=0 を用いる。
  • posterior predictor にテキストと話者情報を注入して、変分後方を真の後方に合わせる。
  • z_H, z_L からなる階層的潜在スキームで容量キャップ C_H と C_L を用い、高次と低次の変動を分離する。
  • 条件付き依存性を持つモデルと持たないモデルで同一テキスト転送、テキスト間スタイル転送、話者間転送、priorサンプリングを比較して評価する。
  • 実験では128次元埋め込みを固定容量設定として用いる。

実験結果

リサーチクエスチョン

  • RQ1埋め込み容量を用いて、韻律/スタイル転送のための変分法とヒューリスティックTTSを比較する統一指標として機能するか。
  • RQ2変分後方をテキストと話者で条件付けることで高忠実度の転送と自然なpriorサンプリングを実現できるか。
  • RQ3階層的潜在分解が転送忠実度とサンプル多様性にどう影響するか。
  • RQ4Lagrangianアプローチによる固定容量が、単一話者および複数話者TTSで堅牢な転送とサンプリングをもたらすか。

主な発見

  • 埋め込み容量の制御は一貫した転送挙動を可能にする:テキストと話者条件付けで容量を高くすると転送が向上し、priorサンプルの品質を損なわない。
  • Var+Txt+Spk は話者間転送をより強く行い、SpkID の指向性とMOSの改善によって対象話者識別を ground truth に近づける。
  • prior からのサンプリングは条件付き依存性を含めた場合に有効で自然な音声を保ち、階層的潜在変数は忠実度とサンプル変動のトレードオフを可能にする(C_H と C_L による)。
  • 話者間転送では Var+Txt+Spk が transferred samples に対して MOS 4.099、SpkID 95.8%、prior-sample MOS 3.906、SpkID 94.9% を達し、ベースラインは MOS ~4.086、SpkID ~95.7%(ground truth は MOS 4.535、SpkID 96.9%)となる。
  • 階層的潜在変数は C_H を上げると参照距離(MCD-DTW)が低下し、C_L を上げるとサンプル間変動が増え、転送の現実感を制御可能にする。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。