Skip to main content
QUICK REVIEW

[論文レビュー] Probabilistic Modelling of Morphologically Rich Languages

Jan A. Botha|arXiv (Cornell University)|Jan 1, 2014
Natural Language Processing Techniques参考文献 177被引用数 4
ひとこと要約

本学位論文は、語彙的構造が複雑な言語における形態素の部分構造を明示的に組み込む確率的言語モデルを提案する。ベイジアンn-gramモデルと分散表現による形態素表現を用いて、スムージングと一般化を向上させる。連続的および不連続的な形態素の教師なし発見を可能にする新しいモデルを導入し、形態素分割および機械翻訳などの下流タスクで優れた性能を示す。

ABSTRACT

This thesis investigates how the sub-structure of words can be accounted for in probabilistic models of language. Such models play an important role in natural language processing tasks such as translation or speech recognition, but often rely on the simplistic assumption that words are opaque symbols. This assumption does not fit morphologically complex language well, where words can have rich internal structure and sub-word elements are shared across distinct word forms. Our approach is to encode basic notions of morphology into the assumptions of three different types of language models, with the intention that leveraging shared sub-word structure can improve model performance and help overcome data sparsity that arises from morphological processes. In the context of n-gram language modelling, we formulate a new Bayesian model that relies on the decomposition of compound words to attain better smoothing, and we develop a new distributed language model that learns vector representations of morphemes and leverages them to link together morphologically related words. In both cases, we show that accounting for word sub-structure improves the models' intrinsic performance and provides benefits when applied to other tasks, including machine translation. We then shift the focus beyond the modelling of word sequences and consider models that automatically learn <em>what</em> the sub-word elements of a given language are, given an unannotated list of words. We formulate a novel model that can learn discontiguous morphemes in addition to the more conventional contiguous morphemes that most previous models are limited to. This approach is demonstrated on Semitic languages, and we find that modelling discontiguous sub-word structures leads to improvements in the task of segmenting words into their contiguous morphemes.

研究の動機と目的

  • 語彙的構造が複雑な言語の言語モデルにおけるデータスパarsityを、語彙的単位の構造を活用することで解消すること。
  • 合成語のベイジアン分解を用いてn-gramモデルのスムージングと一般化を向上させること。
  • 形態素のベクトル表現を学習する分散言語モデルを開発し、関連語形の類似性をモデル化すること。
  • 教師なしで、連続的および不連続的な形態素単位を、アノテーションのない語彙リストから発見できること。
  • 形態素分割および機械翻訳などの下流NLPタスクに、学習した形態素表現を統合すること。

提案手法

  • 合成語を形態素に分解するベイジアンn-gramモデルを提案し、スムージングを向上させ、データスパarsityに対処する。
  • 密なベクトル表現を学習する分散言語モデルを設計し、それらを用いて語形の類似性をモデル化する。
  • 不連続形態素を明示的に扱える、教師なし形態素発見のための新しい確率的モデルを定式化する。
  • 形態的複雑性が高く、不連続形態素が一般的なセムitic言語にモデルを適用する。
  • 変分推論と非パrametricベイジアン手法を用いて、モデルのパラメータを推定し、形態的構造を推論する。
  • 学習した形態素表現を、機械翻訳や形態素分割などの下流NLPタスクに統合する。

実験結果

リサーチクエスチョン

  • RQ1語の部分構造をモデリングすることで、内在的言語モデル性能とデータ効率が向上するか?
  • RQ2分散表現による形態素表現が、形態的に関連する語の間で一般化をどの程度向上させるか?
  • RQ3確率的モデルは、語彙的複雑な言語において、アノテーションのない語彙リストから不連続形態素を発見できるか?
  • RQ4部分構造モデリングは、形態素分割および機械翻訳タスクの性能にどのように影響を与えるか?
  • RQ5形態的分解を組み込むことで、n-gramモデルにおけるスムージングが向上し、パープレキシティが低下するか?

主な発見

  • 形態的分解を施したベイジアンn-gramモデルは、標準n-gramモデルと比較して、より良好なスムージングと低いパープレキシティを達成する。
  • 学習済み形態素ベクトルを用いた分散言語モデルは、形態的豊富なテストセットにおいて一般化性能が向上し、低いパープレキシティを達成する。
  • 提案された教師なしモデルは、セムイト的言語において連続的および不連続的形態素を効果的に同定でき、連続的単位に制限されたモデルを上回る性能を示す。
  • 発見された形態素構造を事前知識または特徴入力として用いることで、形態素分割の性能が向上する。
  • 形態的インフォームドな言語モデルを統合することで、ニューラル機械翻訳システムに測定可能な改善効果が得られる。
  • モデルは、特にアグルーティブ言語およびファズィブ言語の多様な形態的パラダイムにわたり、頑健性を示す。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。