[論文レビュー] LRSpeech: Extremely Low-Resource Speech Synthesis and Recognition
LRSpeechは、事前学習、TTSとASRの間での二重変換、知識蒸留を用いて、極めて低資源な言語における音声合成(TTS)および自動音声認識(ASR)のための低データソリューションを提案する。1.29時間のペアドデータでのリトゥアニア語において、TTSで98.6%の聞き取りやすさと3.65のMOS、ASRで17.04%のWERを達成し、希少言語における産業的実用性を示している。
Speech synthesis (text to speech, TTS) and recognition (automatic speech recognition, ASR) are important speech tasks, and require a large amount of text and speech pairs for model training. However, there are more than 6,000 languages in the world and most languages are lack of speech training data, which poses significant challenges when building TTS and ASR systems for extremely low-resource languages. In this paper, we develop LRSpeech, a TTS and ASR system under the extremely low-resource setting, which can support rare languages with low data cost. LRSpeech consists of three key techniques: 1) pre-training on rich-resource languages and fine-tuning on low-resource languages; 2) dual transformation between TTS and ASR to iteratively boost the accuracy of each other; 3) knowledge distillation to customize the TTS model on a high-quality target-speaker voice and improve the ASR model on multiple voices. We conduct experiments on an experimental language (English) and a truly low-resource language (Lithuanian) to verify the effectiveness of LRSpeech. Experimental results show that LRSpeech 1) achieves high quality for TTS in terms of both intelligibility (more than 98% intelligibility rate) and naturalness (above 3.5 mean opinion score (MOS)) of the synthesized speech, which satisfy the requirements for industrial deployment, 2) achieves promising recognition accuracy for ASR, and 3) last but not least, uses extremely low-resource training data. We also conduct comprehensive analyses on LRSpeech with different amounts of data resources, and provide valuable insights and guidances for industrial deployment. We are currently deploying LRSpeech into a commercialized cloud speech service to support TTS on more rare languages.
研究の動機と目的
- TTSおよびASRシステムにおける6,000種以上の低リソース・希少言語の音声学習データの不足に対処すること。
- 極めて低リソースな言語に対し、スケーラブルで低コストなソリューションを開発し、産業的展開を可能とすること。
- 最小限のペアドデータに加え、非ペアド音声およびクロスタスク知識を活用してTTSおよびASRのパフォーマンスを向上させること。
- 希少言語のTTSおよびASRサービスを商業クラウドプラットフォームに展開可能とする。
提案手法
- 豊富なリソースを持つ言語で事前学習し、その後低リソース言語で微調整することで、言語的および音声的知識を転送する。
- TTSとASRの間での二重変換により、各モデルが偽ペアドデータを生成し、相互に段階的に改善する。
- TTSをターゲットスピーカーの声に適応させ、複数の発話者に対してASRの耐性を高めるために知識蒸留を適用する。
- 確認済みおよび未確認の発話者からの非ペアド音声データを活用して、モデルの一般化性能を向上させる。
- 非ペアドテキストと合成音声を用いて、自己教師学習によりモデル性能をさらに向上させる。
- マルチスケーラーの低品質非ペアド音声と最小限のペアドデータを組み合わせ、データ効率を最大化する。
実験結果
リサーチクエスチョン
- RQ1低リソース環境下で、数分間のペアドデータのみでTTSおよびASRモデルが高い性能を達成できるか?
- RQ2最小限の監視情報のもとで、TTSとASRの間の二重変換が相互の性能向上にどの程度効果的か?
- RQ3合成音声からの知識蒸留は、低リソース言語におけるASRの正確性をどの程度向上できるか?
- RQ4確認済みおよび未確認の発話者からの非ペアド音声の統合が、モデルの一般化に与える影響は?
- RQ5提案されたフレームワークは、希少言語のサポートを目的とした商業クラウドサービスに展開可能か?
主な発見
- LRSpeechはリトゥアニア語において98.60%の聞き取りやすさと3.65の平均意見スコア(MOS)を達成し、産業的展開基準を満たしている。
- ASRモデルは、ペアド学習データがたった1.29時間のリトゥアニア語において、語誤り率(WER)17.04%、文字誤り率(CER)10.30%を達成した。
- 低品質ペアドデータの量を増やすことでTTSのパフォーマンスが向上し、データ量が多いほど精度が向上する。
- 確認済みおよび未確認の発話者からの非ペアド音声を統合することで、TTSおよびASRの精度が顕著に向上した。
- 合成音声データを用いた知識蒸留によりASR性能が向上し、合成データの割合が高いほどより良い結果が得られた。
- 本システムは現在、希少言語のサポートを目的とした商業クラウドTTSサービスへの展開が進行中である。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。