Skip to main content
QUICK REVIEW

[論文レビュー] Analyzing and Improving Statistical Language Models for Speech Recognition

Ueberla, Joerg P.|ArXiv.org|Jun 17, 1994
Speech Recognition and Synthesis参考文献 13被引用数 11
ひとこと要約

本稿では、パーセプレキシティに密接に関連する指標である総確率の対数(LTP)を用いて、音声認識における統計的言語モデルの弱みを形式的に分析する手法を提案する。未知語の処理や長距離依存関係の扱いにおけるモデルの欠陥を特定することで、14–21%の性能向上を達成する改良型バイポスモデルと、拡張された言語的知識を統合する一般化されたNポスモデルを提案し、複雑な認識タスクにおける耐障害性を向上させる。

ABSTRACT

In many current speech recognizers, a statistical language model is used to indicate how likely it is that a certain word will be spoken next, given the words recognized so far. How can statistical language models be improved so that more complex speech recognition tasks can be tackled? Since the knowledge of the weaknesses of any theory often makes improving the theory easier, the central idea of this thesis is to analyze the weaknesses of existing statistical language models in order to subsequently improve them. To that end, we formally define a weakness of a statistical language model in terms of the logarithm of the total probability, LTP, a term closely related to the standard perplexity measure used to evaluate statistical language models. We apply our definition of a weakness to a frequently used statistical language model, called a bi-pos model. This results, for example, in a new modeling of unknown words which improves the performance of the model by 14% to 21%. Moreover, one of the identified weaknesses has prompted the development of our generalized N-pos language model, which is also outlined in this thesis. It can incorporate linguistic knowledge even if it extends over many words and this is not feasible in a traditional N-pos model. This leads to a discussion of whatknowledge should be added to statistical language models in general and we give criteria for selecting potentially useful knowledge. These results show the usefulness of both our definition of a weakness and of performing an analysis of weaknesses of statistical language models in general.

研究の動機と目的

  • 音声認識で用いられる既存の統計的言語モデルにおける弱みを特定し、形式的に定式化すること。
  • パーセプレキシティに密接に関連する形式的指標、すなわち総確率の対数(LTP)を用いて、これらの弱みを体系的に分析・是正することでモデル性能を向上させること。
  • 複数語にわたる言語的知識を統合できる一般化されたNポス言語モデルを構築すること。
  • 統計的言語モデルに統合するのに適した言語的知識の選定基準を確立すること。
  • モデルの弱みを分析することで、音声認識の正確性に測定可能な改善がもたらされることを示すこと。

提案手法

  • パーセプレキシティに密接に関連する形式的指標である総確率の対数(LTP)を用いて、モデルの弱みを定式化すること。
  • LTPに基づく弱み分析を、音声認識分野で広く用いられているバイポス言語モデルに適用すること。
  • 特定された弱みに基づいて、新たな未知語モデル化戦略を設計し、耐障害性と性能を向上させること。
  • 局所的なn-gramを超えて、より長い語列にわたる言語的知識を捉えることができる一般化されたNポス言語モデルを導入すること。
  • LTP指標を用いてモデルのバリエーションを評価・比較し、改善の客観的評価を保証すること。
  • 統計的言語モデルに統合可能な有望な言語的知識の選定・統合のための基準を確立すること。

実験結果

リサーチクエスチョン

  • RQ1音声認識用統計的言語モデルの弱みは、どのように形式的に定義・測定できるか?
  • RQ2バイポスモデルに見られる具体的なモデル欠陥は何か。それらはどのように是正可能か?
  • RQ3一般化されたNポスモデルは、従来のn-gramモデルが欠いているように、複数語にわたる言語的知識を効果的に統合できるか?
  • RQ4統計的言語モデルに統合するにあたり、有用な言語的知識を選定するにあたっての基準は何であるべきか?
  • RQ5モデルの弱みを分析することで、認識正確性に測定可能な改善がどの程度達成できるか?

主な発見

  • 提案された未知語モデル化戦略により、バイポスモデルの性能が14%から21%向上した。
  • 一般化されたNポス言語モデルは、標準的なn-gramモデルに欠如している複数語にわたる言語的知識の統合に成功した。
  • LTP指標は、モデルの弱みを効果的に特定し、言語モデル設計における的を射た改善を導くのに有効であった。
  • 分析により、従来のn-gramモデルが長距離の言語的依存関係を捉えられていないことが明らかになり、一般化されたNポスフレームワークの導入が正当化された。
  • 本研究では、形式的な弱み分析が、体系的かつ測定可能な言語モデル性能の向上をもたらすことを示した。
  • 有用な言語的知識の選定フレームワークは、統計的言語モデルを洗練させるための原則的かつ整合性のあるアプローチを提供した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。