Skip to main content
QUICK REVIEW

[論文レビュー] Leveraging Deep Neural Network Activation Entropy to cope with Unseen Data in Speech Recognition

Vikramjit Mitra, Horacio Franco|arXiv (Cornell University)|Aug 31, 2017
Speech Recognition and Synthesis参考文献 25被引用数 4
ひとこと要約

本稿では、自動音声認識における未知のデータ条件を同定し、それに適応するための信頼性指標として、深層ニューラルネットワークの活性化エントロピーを提案する。隠れ層におけるニューロン活性化の短時間ウィンドウでのランニングエントロピーを計算することで、信頼性の高いデータを特定し、無教師モデル適応に用いる。これにより、ベンチマーク音声認識タスクにおける未知の音響条件下での性能低下が顕著に軽減される。

ABSTRACT

Unseen data conditions can inflict serious performance degradation on systems relying on supervised machine learning algorithms. Because data can often be unseen, and because traditional machine learning algorithms are trained in a supervised manner, unsupervised adaptation techniques must be used to adapt the model to the unseen data conditions. However, unsupervised adaptation is often challenging, as one must generate some hypothesis given a model and then use that hypothesis to bootstrap the model to the unseen data conditions. Unfortunately, reliability of such hypotheses is often poor, given the mismatch between the training and testing datasets. In such cases, a model hypothesis confidence measure enables performing data selection for the model adaptation. Underlying this approach is the fact that for unseen data conditions, data variability is introduced to the model, which the model propagates to its output decision, impacting decision reliability. In a fully connected network, this data variability is propagated as distortions from one layer to the next. This work aims to estimate the propagation of such distortion in the form of network activation entropy, which is measured over a short- time running window on the activation from each neuron of a given hidden layer, and these measurements are then used to compute summary entropy. This work demonstrates that such an entropy measure can help to select data for unsupervised model adaptation, resulting in performance gains in speech recognition tasks. Results from standard benchmark speech recognition tasks show that the proposed approach can alleviate the performance degradation experienced under unseen data conditions by iteratively adapting the model to the unseen datas acoustic condition.

研究の動機と目的

  • 分布シフトによって生じる未知のデータ条件下での音声認識システムの性能低下を是正すること。
  • 無教師適応の場面において、モデル仮説の信頼性指標を確立すること。
  • 活性化エントロピーを用いてDNN内での歪み伝搬を推定し、より良いデータ選択を実現すること。
  • エントロピーに基づくデータ選択が、未知の音響条件下でのモデル適応性能を向上させることを実証すること。
  • ラベル付きデータが未知ドメインから得られない状況でも、実用的かつ計算効率の良いロバスト音声認識手法を提供すること。

提案手法

  • 全結合DNNの各隠れ層からのニューロン出力に対して、短時間ウィンドウでの活性化エントロピーを測定する。
  • 特定の層内の全ニューロンに対する要約エントロピーを計算し、ネットワーク応答の不確実性を定量化する。
  • 得られたエントロピースコアを信頼性指標として用い、高不確実性(未知に似た)入力を同定する。
  • 高エントロピーのサンプルを無教師モデル適応用に選択し、ロバスト性を向上させる。
  • 選択されたデータを用いて反復的適応を適用し、未知の条件下での音声モデルを精緻化する。
  • 標準的な音声認識ベンチマークを用いて、手法の有効性を検証する。

実験結果

リサーチクエスチョン

  • RQ1活性化エントロピーは、音声認識における未知データの同定に信頼性のある信頼性指標として機能するか?
  • RQ2エントロピーに基づくデータ選択は、無教師モデル適応をどの程度効果的に向上させるか?
  • RQ3エントロピー駆動の適応は、未知の音響条件下での性能低下を軽減するか?
  • RQ4活性化エントロピーは、DNN内での入力変動に起因する歪み伝搬を捉えることができるか?
  • RQ5エントロピーに基づくデータ選択は、標準ベンチマークにおける認識正答率にどのような影響を与えるか?

主な発見

  • 活性化エントロピーは、入力歪みに起因する不確実性の伝搬を捉えることで、未知のデータ条件からの入力を効果的に同定する。
  • エントロピーに基づくデータ選択手法は、標準音声認識ベンチマークにおける未知の音響条件下でのモデル適応性能を向上させる。
  • 高エントロピーのサンプルを用いた反復的適応により、未知のデータ条件下での性能低下が顕著に軽減される。
  • ベースラインの適応戦略と比較して、未知のテストセットにおける単語誤り率(WER)の低減が明確に測定される。
  • 本手法は計算効率が高く、未知ドメインからの追加ラベル付きデータを必要としない。
  • 結果から、分布シフト下における意思決定の信頼性とDNN活性化におけるエントロピーの相関が強く示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。