Skip to main content
QUICK REVIEW

[論文レビュー] Improving the Performance of English-Tamil Statistical Machine Translation System using Source-Side Pre-Processing

M. Anand Kumar, V. Dhanalakshmi|arXiv (Cornell University)|Sep 29, 2014
Natural Language Processing Techniques参考文献 14被引用数 7
ひとこと要約

この論文は、品詞タギングや語形還元を含む言語的特徴を統合することで、翻訳精度を向上させるために、英語=タミル語統計的機械翻訳(SMT)システムにおけるソース側前処理を提案する。翻訳の前段階で語彙的および文法的情報を強化することで、BLEUスコアが著しく向上し、限られた並列コーパスでも言語的前処理が性能向上に寄与することを示している。

ABSTRACT

Machine Translation is one of the major oldest and the most active research area in Natural Language Processing. Currently, Statistical Machine Translation (SMT) dominates the Machine Translation research. Statistical Machine Translation is an approach to Machine Translation which uses models to learn translation patterns directly from data, and generalize them to translate a new unseen text. The SMT approach is largely language independent, i.e. the models can be applied to any language pair. Statistical Machine Translation (SMT) attempts to generate translations using statistical methods based on bilingual text corpora. Where such corpora are available, excellent results can be attained translating similar texts, but such corpora are still not available for many language pairs. Statistical Machine Translation systems, in general, have difficulty in handling the morphology on the source or the target side especially for morphologically rich languages. Errors in morphology or syntax in the target language can have severe consequences on meaning of the sentence. They change the grammatical function of words or the understanding of the sentence through the incorrect tense information in verb. Baseline SMT also known as Phrase Based Statistical Machine Translation (PBSMT) system does not use any linguistic information and it only operates on surface word form. Recent researches shown that adding linguistic information helps to improve the accuracy of the translation with less amount of bilingual corpora. Adding linguistic information can be done using the Factored Statistical Machine Translation system through pre-processing steps. This paper investigates about how English side pre-processing is used to improve the accuracy of English-Tamil SMT system.

研究の動機と目的

  • 統計的機械翻訳における低資源言語対(例:英語=タミル語)の課題に対処すること。
  • 広範な並列コーパスを必要とせずに、ソース側における言語的前処理が翻訳品質を向上させることを調査すること。
  • 語彙的および文法的特徴が、タミル語のような語彙的豊富なターゲット言語におけるSMTパフォーマンスに与える影響を調査すること。
  • 前処理済みソース側言語的情報を利用した要因SMTの有効性を評価すること。
  • 前処理を通じて、言語に依存しない低資源に配慮したSMTの改善手法を提供すること。

提案手法

  • 英語ソース文に品詞(POS)タギングを適用し、文法的役割を特定する。
  • 英語語彙に対して語形還元を実施し、屈折変化のばらつきを低減し、語形を標準化する。
  • POSおよび語形情報を要因SMTフレームワークに統合し、翻訳モデリングを支援する。
  • 前処理済みソース側特徴をSMTシステムの追加要因として統合し、アライメントおよび翻訳意思決定を改善する。
  • 言語的前処理を施した限られた並列コーパスを用いて、フレーズベースSMTシステムを学習する。
  • BLEUスコアなどの標準的メトリクスを用いて翻訳品質を評価し、改善度を測定する。

実験結果

リサーチクエスチョン

  • RQ1低資源の英語=タミル語SMTにおいて、ソース側言語的前処理が翻訳品質を向上させ得るか?
  • RQ2品詞タギングと語形還元は、英語=タミル語翻訳の要因SMTシステムのパフォーマンスにどのように影響するか?
  • RQ3語彙的および構文的課題を克服するため、前処理がタミル語のような語彙的豊富な言語への翻訳においてどの程度効果を発揮するか?
  • RQ4言語的前処理は、SMTシステムにおける大規模並列コーパスへの依存度を低下させるか?
  • RQ5前処理済みソース側特徴を用いた場合、翻訳品質の向上が統計的に有意であるか?

主な発見

  • ソース側前処理の導入により、英語=タミル語SMTシステムのBLEUスコアが顕著に向上した。
  • 語形還元と品詞タギングにより、より良い語のアライメントが実現し、翻訳のあいまいさが低減した。
  • 前処理済み特徴を用いた要因SMTシステムは、言語的情報を含まないベースラインPBSMTシステムを上回った。
  • 動詞の時制や一致の誤りの処理において、特に顕著な改善が見られた。
  • 限られた並列データでも、言語的前処理がSMTパフォーマンスを向上させることを確認した。
  • 語彙的豊富なターゲット言語における低資源環境でも、本手法が実用的かつ有効であることを示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。