Skip to main content
QUICK REVIEW

[論文レビュー] Multi-domain machine translation enhancements by parallel data extraction from comparable corpora

Krzysztof Wołk, Emilia Rejmund|arXiv (Cornell University)|Mar 22, 2016
Natural Language Processing Techniques参考文献 2被引用数 7
ひとこと要約

本稿では、Wikipedia やドメイン固有のウェブコレクションなどの類似コーパスから、自動ウェブクローリングと類似度ベースのフィルタリングを用いて、高品質な並列文を抽出する非教師ありでスケーラブルな手法を提案する。このアプローチにより、事前にある並列リソースが不要な状況でも、ドメイン適応型の並列データを生成することで、多様なドメイン(例:医学、映画、議会会議)における統計的機械翻訳の性能が顕著に向上し、BLEUスコアで測定可能な向上が得られた。

ABSTRACT

Parallel texts are a relatively rare language resource, however, they constitute a very useful research material with a wide range of applications. This study presents and analyses new methodologies we developed for obtaining such data from previously built comparable corpora. The methodologies are automatic and unsupervised which makes them good for large scale research. The task is highly practical as non-parallel multilingual data occur much more frequently than parallel corpora and accessing them is easy, although parallel sentences are a considerably more useful resource. In this study, we propose a method of automatic web crawling in order to build topic-aligned comparable corpora, e.g. based on the Wikipedia or Euronews.com. We also developed new methods of obtaining parallel sentences from comparable data and proposed methods of filtration of corpora capable of selecting inconsistent or only partially equivalent translations. Our methods are easily scalable to other languages. Evaluation of the quality of the created corpora was performed by analysing the impact of their use on statistical machine translation systems. Experiments were presented on the basis of the Polish-English language pair for texts from different domains, i.e. lectures, phrasebooks, film dialogues, European Parliament proceedings and texts contained medicines leaflets. We also tested a second method of creating parallel corpora based on data from comparable corpora which allows for automatically expanding the existing corpus of sentences about a given domain on the basis of analogies found between them. It does not require, therefore, having past parallel resources in order to train a classifier.

研究の動機と目的

  • 低リソースおよびドメイン特化翻訳タスクにおける並列コーパスの不足を解決すること。
  • 大規模な類似コーパスから並列文を抽出する非教師ありでスケーラブルな手法を開発すること。
  • 非並列ソースから抽出したドメイン特化並列データで統計的機械翻訳システムを強化し、性能を向上させること。
  • 事前にある並列トレーニングデータが不要な状況でも、類似性に基づく文マッチングを用いて既存の並列コーパスを自動で拡張できること。
  • 抽出データの影響が複数のドメインおよび言語対(特にポーランド語-英語)における翻訳品質に与える影響を評価すること。

提案手法

  • Wikipedia などのマルチリンガルソースからトピックが一致する類似コーパスを構築するための自動ウェブクローリング。
  • 類似コーパス内の候補となる並列文を特定するための非教師あり文のアライメント技術の適用。
  • 不一致または部分的に同等の翻訳を除去する類似度ベースのフィルタリングを用いて、データ品質を向上させること。
  • 言語間の文埋め込みモデルを活用し、言語間での文レベルの対応関係を検出すること。
  • 異なるドメイン間で類似した文構造を特定することで、既存の並列コーパスを拡張する手法の設計。
  • 抽出された並列データを統計的機械翻訳システムに統合し、ドメイン特化のための適応を図ること。

実験結果

リサーチクエスチョン

  • RQ1事前にある並列データが不要な状況でも、類似コーパスから並列文を効果的に抽出できるか?
  • RQ2抽出された並列データの品質は、手動でキュレートされたコーパスと比較して翻訳性能においてどの程度優れているか?
  • RQ3提案手法が医療、映画、議会会議など多様なドメインにおける機械翻訳性能をどの程度向上できるか?
  • RQ4ポーランド語-英語に限らず、他の言語対にも一般化可能か?
  • RQ5類似性に基づく拡張手法は、類似データを用いて既存の並列コーパスをどの程度効果的に拡張できるか?

主な発見

  • 提案手法は、類似コーパスから高品質な並列文を効果的に抽出でき、複数のドメインにおける翻訳性能を顕著に向上させた。
  • 抽出データの活用により、レクチャーやフレーズブック、映画台詞、医療用添付文書など、すべてのテストドメインでBLEUスコアに測定可能な向上が得られた。
  • フィルタリング機構により、不一致や部分的に同等の翻訳が効果的に削除され、データの信頼性が向上した。
  • 類似性に基づく拡張手法により、事前の並列データが不要な状況でもコーパスの拡大が可能であり、スケーラビリティと適応性が実証された。
  • 異なる言語間でも本手法は効果的かつスケーラブルであり、低リソースおよびドメイン特化環境で顕著な性能向上が観察された。
  • 評価により、抽出データがランダムまたはフィルタリングなしの類似データで学習したベースラインシステムを上回ることが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。