Skip to main content
QUICK REVIEW

[論文レビュー] FluentNet: End-to-End Detection of Speech Disfluency with Deep Learning

Tedd Kourkounakis, Amirhossein Hajavi|arXiv (Cornell University)|Sep 23, 2020
Stuttering Research and Treatment参考文献 46被引用数 9
ひとこと要約

FluentNet は、スペクトル特徴の学習に Squeeze-and-Excitation 残差 CNN を、時間的モデリングに双方向 LSTMs を、グローバルなアテンションメカニズムを用いたエンド・ツー・エンドのディーブラーニングモデルであり、音声の不順応(音声・語・フレーズの繰り返し、訂正、挿入語、長音化)の6種類を検出する。UCLASS データセットで最先端の性能を達成し、LibriStutter と呼ばれる新しい合成データセットに対しても良好な汎化性能を示す。

ABSTRACT

Strong presentation skills are valuable and sought-after in workplace and classroom environments alike. Of the possible improvements to vocal presentations, disfluencies and stutters in particular remain one of the most common and prominent factors of someone's demonstration. Millions of people are affected by stuttering and other speech disfluencies, with the majority of the world having experienced mild stutters while communicating under stressful conditions. While there has been much research in the field of automatic speech recognition and language models, there lacks the sufficient body of work when it comes to disfluency detection and recognition. To this end, we propose an end-to-end deep neural network, FluentNet, capable of detecting a number of different disfluency types. FluentNet consists of a Squeeze-and-Excitation Residual convolutional neural network which facilitate the learning of strong spectral frame-level representations, followed by a set of bidirectional long short-term memory layers that aid in learning effective temporal relationships. Lastly, FluentNet uses an attention mechanism to focus on the important parts of speech to obtain a better performance. We perform a number of different experiments, comparisons, and ablation studies to evaluate our model. Our model achieves state-of-the-art results by outperforming other solutions in the field on the publicly available UCLASS dataset. Additionally, we present LibriStutter: a disfluency dataset based on the public LibriSpeech dataset with synthesized stutters. We also evaluate FluentNet on this dataset, showing the strong performance of our model versus a number of benchmark techniques.

研究の動機と目的

  • 多様な音声不順応タイプの正確な検出を目的としたエンド・ツー・エンドのディーブラーニングモデルの開発。
  • LibriSpeech を基盤として、実際の不順応データの不足を補うために合成データセットである LibriStutter を作成することによる、ラベル付き不順応データの不足への対応。
  • 単なるフィラー語の検出を越えて、複雑なスティッターパターンをモデル化することによる不順応検出の向上。
  • Squeeze-and-Excitation やアテンションメカニズムといったキーコンポーネントが不順応認識に与える寄与の評価。
  • UCLASS や新たに導入された LibriStutter のような公開ベンチマークで最先端の性能を示すこと。

提案手法

  • FluentNet は、生の音声からスペクトルフレームレベルの表現を学習するために、Squeeze-and-Excitation (SE) 残差畳み込みニューラルネットワークを用いる。
  • スティッタード音声における長距離の時間的依存関係をモデル化するために、双方向長短期記憶(BLSTM)層を採用する。
  • BLSTM 層の後にグローバルアテンションメカニズムを適用し、不順応分類のための顕著な音声セグメントに注目する。
  • 学習率 $10^{-4}$ を用いた交差エントロピー損失関数に基づき、Adam を用いてエンド・ツー・エンドで最適化して訓練する。
  • UCLASS データセット上でハイパーパramータのアブレーションを実施し、8つの残差ブロックと2つの BLSTM 層を採用したアーキテクチャを選定した。
  • 合成不順応データセットである LibriStutter は、クリアな LibriSpeech 音声サンプルに合成されたスティッタを挿入することで作成された。

実験結果

リサーチクエスチョン

  • RQ1エンド・ツー・エンドのディーブラーニングモデルは、多様な種類の音声不順応を検出する上で最先端の性能を達成できるか?
  • RQ2Squeeze-and-Excitation やアテンションメカニズムといったコンポーネントは、不順応検出性能にどのように寄与するか?
  • RQ3本モデルは、現実のスティッターパターンを模倣する合成不順応データに対して、どの程度汎化性能を示すか?
  • RQ4本モデルは、訂正や長音化といったレアまたは複雑なパターンを含む多様な不順応タイプに対して、どの程度の性能を示すか?
  • RQ5実データが限られる状況において、LibriStutter のような合成データセットが、不順応検出モデルのトレーニングの代替として有効に機能するか?

主な発見

  • FluentNet は、UCLASS データセットで最先端の性能を達成し、既存の手法を上回る不順応分類性能を示した。
  • UCLASS データセットにおいて、すべての不順応タイプで約 20 エポックでほぼ完全な訓練精度に達した。
  • アブレーションスタディの結果、Squeeze-and-Excitation 要素を削除すると、すべての不順応タイプにおいて最も顕著な精度低下と誤検出率の上昇が観察された。
  • アテンションメカニズムは性能向上に顕著な寄与を示しており、その削除により分類精度が顕著に低下した。
  • モデルは合成された LibriStutter データセットに対しても効果的に汎化しており、性能低下は観察されたが、依然としてベンチマークを上回った。
  • アブレーション結果は UCLASS および LibriStutter の両方で一貫しており、合成データセットの妥当性と現実性を裏付けた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。