Skip to main content
QUICK REVIEW

[论文解读] The Effect of Domain and Diacritics in Yorùbá-English Neural Machine Translation

David Ifeoluwa Adelani, Dana Ruiter|arXiv (Cornell University)|Mar 15, 2021
Natural Language Processing Techniques参考文献 39被引用 6
一句话总结

本文介紹了 MENYO-20k,一個高品質、多領域的英語–約魯巴語平行語料庫,具備標準化的訓練-測試分割,並評估其對神經機器翻譯的影響。透過納入正確的變音符號與領域特定的資料,作者證明微調模型在約魯巴語–英語翻譯中,分別比 Google 和 Facebook 的 M2M 多語模型高出 +8.7 和 +9.1 BLEU。

ABSTRACT

Massively multilingual machine translation (MT) has shown impressive capabilities, including zero and few-shot translation between low-resource language pairs. However, these models are often evaluated on high-resource languages with the assumption that they generalize to low-resource ones. The difficulty of evaluating MT models on low-resource pairs is often due to lack of standardized evaluation datasets. In this paper, we present MENYO-20k, the first multi-domain parallel corpus with a special focus on clean orthography for Yorùbá--English with standardized train-test splits for benchmarking. We provide several neural MT benchmarks and compare them to the performance of popular pre-trained (massively multilingual) MT models both for the heterogeneous test set and its subdomains. Since these pre-trained models use huge amounts of data with uncertain quality, we also analyze the effect of diacritics, a major characteristic of Yorùbá, in the training data. We investigate how and when this training condition affects the final quality and intelligibility of a translation. Our models outperform massively multilingual models such as Google ($+8.7$ BLEU) and Facebook M2M ($+9.1$ BLEU) when translating to Yorùbá, setting a high quality benchmark for future research.

研究动机与目标

  • 解決低資源約魯巴語–英語機器翻譯缺乏標準化、高品質平行語料庫的問題。
  • 探討訓練資料中領域多樣性與變音符號準確性對翻譯品質的影響。
  • 利用精心整理的多領域資料集,建立約魯巴語–英語神經機器翻譯的基準。
  • 評估監督式、半監督式與微調模型在低資源約魯巴語–英語翻譯中的表現。
  • 分析變音符號如何影響翻譯品質,特別是在英語–約魯巴語方向,並評估變音符號還原的效益。

提出的方法

  • 建立 MENYO-20k,一個包含 20,000 筆句子的平行語料庫,涵蓋新聞、TED 演講、廣播、科學/科技與對話文本,由專業翻譯員精心整理。
  • 定義標準化的訓練-開發-測試分割,以支援約魯巴語–英語神經機器翻譯的可重現基準測試。
  • 在 MENYO-20k 併合 JW300 與《聖經》資料集的基礎上,訓練多種監督式、半監督式與微調的神經機器翻譯模型。
  • 使用自動指標(BLEU)與人工評估(流暢度與可理解度)來評估模型表現。
  • 與 Google 的 MNMT 和 Facebook 的 M2M-100 等大型多語模型進行對比。
  • 透過在具備與不具變音符號的資料版本上訓練模型,分析變音符號的影響。
Figure 1: Top: Perplexities of KenLM 5-gram language model learned on different training corpora and tested on subsets of MENYO-20k for English (left) and Yorùbá (right) respectively. Bottom: Vocabulary coverage (%) of different subsets of the MENYO-20k test set per training sets for English (left)
Figure 1: Top: Perplexities of KenLM 5-gram language model learned on different training corpora and tested on subsets of MENYO-20k for English (left) and Yorùbá (right) respectively. Bottom: Vocabulary coverage (%) of different subsets of the MENYO-20k test set per training sets for English (left)

实验结果

研究问题

  • RQ1訓練資料中的領域多樣性如何影響低資源約魯巴語–英語神經機器翻譯模型的泛化能力?
  • RQ2訓練資料中正確的變音符號在多大程度上能提升翻譯品質,特別是在英語–約魯巴語方向?
  • RQ3在 MENYO-20k 上微調的模型是否能超越 Google 的 MNMT 與 Facebook 的 M2M 模型在約魯巴語–英語翻譯中的表現?
  • RQ4訓練資料中存在變音符號如何影響人工評估中的流暢度與可理解度?
  • RQ5當變音符號不一致時,訓練與測試資料之間的領域不匹配對翻譯品質有何影響?

主要发现

  • 在 MENYO-20k 上微調的模型,其在約魯巴語–英語翻譯中,分別比 Google 的多語模型高出 +8.7 BLEU,比 Facebook 的 M2M 模型高出 +9.1 BLEU。
  • 人工評估顯示,即使內容可理解,使用低品質、無變音符號資料訓練的模型(如 Google MNMT)仍顯得較不流暢。
  • 訓練資料中正確的變音符號顯著提升了英語–約魯巴語方向的翻譯品質,特別是在領域不匹配的情況下。
  • 即使僅使用 10,000 筆平行句子,結合 JW300 與《聖經》資料集的 MENYO-20k 在所有領域中均提升了翻譯表現。
  • 本研究證明,在低資源環境下,具備正確變音符號的高品質、語言專用資料,表現優於龐大但雜訊多的多語語料庫。
  • 自動變音符號還原被證明能有效減輕訓練資料中雜訊或遺漏變音符號的負面影響。
The Effect of Domain and Diacritics in Yorùbá-English Neural Machine Translation

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。