Skip to main content
QUICK REVIEW

[論文レビュー] Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents

Priyank Dubey, Bilal Shah|arXiv (Cornell University)|Apr 3, 2022
Speech Recognition and Synthesis被引用数 14
ひとこと要約

本論文では、インド英語の発音を対象とした、転移学習とデータ拡張を用いた、DeepSpeechベースのエンドツーエンドASRシステムの微調整を提案する。Indic TTSデータセットを用いて、事前学習済みのDeepSpeech-0.9.3モデルをインド英語の音声に適応させることで、未訓練モデルや商用サービスよりも顕著な向上を達成し、インドの多様な地方英語発音に対しても優れた性能を示した。

ABSTRACT

Automated Speech Recognition (ASR) is an interdisciplinary application of computer science and linguistics that enable us to derive the transcription from the uttered speech waveform. It finds several applications in Military like High-performance fighter aircraft, helicopters, air-traffic controller. Other than military speech recognition is used in healthcare, persons with disabilities and many more. ASR has been an active research area. Several models and algorithms for speech to text (STT) have been proposed. One of the most recent is Mozilla Deep Speech, it is based on the Deep Speech research paper by Baidu. Deep Speech is a state-of-art speech recognition system is developed using end-to-end deep learning, it is trained using well-optimized Recurrent Neural Network (RNN) training system utilizing multiple Graphical Processing Units (GPUs). This training is mostly done using American-English accent datasets, which results in poor generalizability to other English accents. India is a land of vast diversity. This can even be seen in the speech, there are several English accents which vary from state to state. In this work, we have used transfer learning approach using most recent Deep Speech model i.e., deepspeech-0.9.3 to develop an end-to-end speech recognition system for Indian-English accents. This work utilizes fine-tuning and data argumentation to further optimize and improve the Deep Speech ASR system. Indic TTS data of Indian-English accents is used for transfer learning and fine-tuning the pre-trained Deep Speech model. A general comparison is made among the untrained model, our trained model and other available speech recognition services for Indian-English Accents.

研究の動機と目的

  • 主にアメリカ英語で訓練された既存のASRシステムがインド英語の発音に一般化できない問題を解決すること。
  • インドの英語発音パターンの言語的多様性に特化した、強固なエンドツーエンドASRシステムの開発。
  • 事前学習済みのDeepSpeech-0.9.3モデルを活用した転移学習により、訓練データと計算リソースの要件を低減すること。
  • インド英語の音声に特化した、微調整とデータ拡張を用いたモデル性能最適化。

提案手法

  • インド英語の発音を含むIndic TTSデータセットを用いて、事前学習済みのDeepSpeech-0.9.3モデルを転移学習により微調整した。
  • データ拡張技術を適用して、訓練データの多様性を高め、モデルの頑健性を向上させた。
  • 接続主義的時系列分類(CTC)損失関数を用いた、深層双方向LSTM(Bi-LSTM)ネットワークに基づくシーケンス・ツー・シーケンスアーキテクチャを採用した。
  • 複数のGPUを用いて訓練することで、収束の加速と最適化効率の向上を図った。
  • 未訓練モデル、提案モデル、および商用ASRサービスとの間で語誤り率(WER)を比較して性能を評価した。

実験結果

リサーチクエスチョン

  • RQ1限定的なドメイン固有データで、事前学習済みのDeepSpeechモデルを用いた転移学習が、インド英語の発音に効果的に適応できるか。
  • RQ2データ拡張は、インド英語発音向けASRモデルの頑健性と精度にどのように影響するか。
  • RQ3微調整されたDeepSpeechモデルは、インド英語発音認識タスクにおいて、未訓練モデルに比べてどの程度の性能向上を達成するか。
  • RQ4提案システムは、既存の商用ASRサービスと比較して、インド英語の発音認識においてどの程度の精度を示すか。

主な発見

  • 微調整されたDeepSpeechモデルは、インド英語発音において、未訓練ベースラインモデルと比較して顕著な語誤り率(WER)の低減を達成した。
  • 提案システムは、インド英語の発音において、商用ASRサービスを上回るWERを示し、ドメイン固有の適応性が優れていることを示した。
  • データ拡張は、インド国内のレアまたはリソースが乏しい地方発音において特に、モデルの一般化性能の向上に寄与した。
  • 転移学習により、追加の訓練データと計算コストを最小限に抑えながら、事前学習済みモデルの効果的な適応が可能となった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。