Skip to main content
QUICK REVIEW

[論文レビュー] Jira: a Kurdish Speech Recognition System Designing and Building Speech Corpus and Pronunciation Lexicon

Hadi Veisi, Hawre Hosseini|arXiv (Cornell University)|Feb 15, 2021
Speech Recognition and Synthesis参考文献 19被引用数 5
ひとこと要約

本稿では、制御された環境とクラウドソーシング環境の両方で576人の話者から収集した43.68時間の音声コーパスと、60,000語の発音語彙を構築することで、中央クルド語用の最初の大規模語彙音声認識システム「Jira」を提示する。Kaldiツールキットを用いて、SGMM音声モデルを用いることで、多様なトピックでは13.9%の語誤り率を達成し、一般トピックでは4.9%まで低下させた。これはクルド語NLPリソース分野における基盤的貢献である。

ABSTRACT

In this paper, we introduce the first large vocabulary speech recognition system (LVSR) for the Central Kurdish language, named Jira. The Kurdish language is an Indo-European language spoken by more than 30 million people in several countries, but due to the lack of speech and text resources, there is no speech recognition system for this language. To fill this gap, we introduce the first speech corpus and pronunciation lexicon for the Kurdish language. Regarding speech corpus, we designed a sentence collection in which the ratio of di-phones in the collection resembles the real data of the Central Kurdish language. The designed sentences are uttered by 576 speakers in a controlled environment with noise-free microphones (called AsoSoft Speech-Office) and in Telegram social network environment using mobile phones (denoted as AsoSoft Speech-Crowdsourcing), resulted in 43.68 hours of speech. Besides, a test set including 11 different document topics is designed and recorded in two corresponding speech conditions (i.e., Office and Crowdsourcing). Furthermore, a 60K pronunciation lexicon is prepared in this research in which we faced several challenges and proposed solutions for them. The Kurdish language has several dialects and sub-dialects that results in many lexical variations. Our methods for script standardization of lexical variations and automatic pronunciation of the lexicon tokens are presented in detail. To setup the recognition engine, we used the Kaldi toolkit. A statistical tri-gram language model that is extracted from the AsoSoft text corpus is used in the system. Several standard recipes including HMM-based models (i.e., mono, tri1, tr2, tri2, tri3), SGMM, and DNN methods are used to generate the acoustic model. These methods are trained with AsoSoft Speech-Office and AsoSoft Speech-Crowdsourcing and a combination of them. The best performance achieved by the SGMM acoustic model which results in 13.9% of the average word error rate (on different document topics) and 4.9% for the general topic.

研究の動機と目的

  • 中央クルド語(3,000万人以上が話す言語)に不足している音声およびテキストリソースを補完すること。
  • 実際の中央クルド語の発音に代表されるディファオン分布を反映する音声コーパスを設計すること。
  • クルド語の方言間の語彙的変異を考慮した標準化された大規模発音語彙を構築すること。
  • 最先端のASR技術を用いて中央クルド語用の大語彙音声認識システム(LVSR)を構築および評価すること。
  • 再利用可能な音声および言語リソースを提供することで、クルド語の基盤的NLPインfraを確立すること。

提案手法

  • 実際の中央クルド語データのディファオン頻度に一致する文の収集を設計し、言語的代表性を確保した。
  • AsoSoft Speech-Office(制御された、ノイズのない環境)およびAsoSoft Speech-Crowdsourcing(Telegram経由のスマートフォン)の2つの環境を用いて、576人の話者から合計43.68時間の音声を収集した。
  • 両方の環境(オフィスおよびクラウドソーシング)で録音された11のドキュメントトピックを含むテストセットを構築し、評価の堅牢性を確保した。
  • クルド語の方言間の表記のばらつきを解消するためのスクリプト標準化技術を用いて、60,000語の発音語彙を構築した。
  • 自動発音モデリングを適用して、語彙エントリの発音表記を自動生成した。
  • 両方の音声収集環境のデータを統合し、Kaldiツールキットを用いて複数の音声モデル(HMMベース:モノ、トライ1〜トライ3;SGMM;DNN)を訓練した。

実験結果

リサーチクエスチョン

  • RQ1環境ノイズを最小限に抑えつつ、現実世界の可変性を反映する大規模で言語的に代表的な中央クルド語音声コーパスをどのように収集できるか?
  • RQ2クルド語の方言間で生じる表記のばらつきを効果的に標準化し、大規模語彙の正確な発音を生成する方法は何か?
  • RQ3制御された環境とクラウドソーシングの両方の音声データを統合することで、クルド語LVSRシステムの性能にどのような影響を与えるか?
  • RQ4語誤り率を低く抑えるために、中央クルド語ASRにおいて最適な音声モデリング手法(例:SGMM、DNN)は何か?
  • RQ5AsoSoftテキストコーパス上で学習された言語モデルは、多様なトピックにおける認識性能をどの程度向上させるか?

主な発見

  • AsoSoft Speech-OfficeおよびAsoSoft Speech-Crowdsourcingの両データセットは、合計で576人の話者から43.68時間の高品質で多様な音声データを収集した。
  • SGMM音声モデルを用いた場合、11の異なるドキュメントトピックの平均で13.9%の語誤り率を達成した。
  • 一般トピックでは語誤り率が4.9%まで低下し、広範な分野にまたがる非ドメイン特化音声において優れた性能を示した。
  • スクリプト標準化と自動発音モデリングを用いて、方言のばらつきの課題を克服した60,000語の発音語彙が正常に構築された。
  • SGMMモデルはHMMおよびDNNベースラインを上回り、低リソースのクルド語ASRにおいて有効性を示した。
  • AsoSoftテキストコーパス上で学習された3-gram言語モデルの統合により、トピック全体にわたる認識の堅牢性が著しく向上した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。