Skip to main content
QUICK REVIEW

[Paper Review] MENYO-20k: A Multi-domain English-Yorùbá Corpus for Machine Translation and Domain Adaptation.

David Ifeoluwa Adelani, Dana Ruiter|arXiv (Cornell University)|Mar 15, 2021
Natural Language Processing Techniques4 citations
TL;DR

This paper introduces MENYO-20k, the first multi-domain parallel corpus for the low-resource Yoruba–English language pair, with standardized train-test splits. It demonstrates that fine-tuning generic multilingual models on MENYO-20k yields significant BLEU gains—+9.9 and +8.6 over Facebook's M2M-100 and Google's multilingual NMT, respectively—outperforming pre-trained models in both zero- and few-shot translation settings.

ABSTRACT

Massively multilingual machine translation (MT) has shown impressive capabilities, including zero and few-shot translation between low-resource language pairs. However, these models are often evaluated on high-resource languages with the assumption that they generalize to low-resource ones. The difficulty of evaluating MT models on low-resource pairs is often due the lack of standardized evaluation datasets. In this paper, we present MENYO-20k, the first multi-domain parallel corpus for the low-resource Yoruba--English (yo--en) language pair with standardized train-test splits for benchmarking. We provide several neural MT (NMT) benchmarks on this dataset and compare to the performance of popular pre-trained (massively multilingual) MT models, showing that, in almost all cases, our simple benchmarks outperform the pre-trained MT models. A major gain of BLEU $+9.9$ and $+8.6$ (en2yo) is achieved in comparison to Facebook's M2M-100 and Google multilingual NMT respectively when we use MENYO-20k to fine-tune generic models.

Motivation & Objective

  • To address the lack of standardized evaluation datasets for low-resource language pairs like Yoruba–English.
  • To enable reliable benchmarking of neural machine translation models on low-resource language pairs.
  • To improve translation performance for Yoruba by providing a curated, multi-domain parallel corpus for fine-tuning.
  • To evaluate the effectiveness of pre-trained multilingual models versus fine-tuned models on a low-resource language pair.
  • To demonstrate that domain-specific fine-tuning on MENYO-20k significantly boosts translation quality over generic pre-trained models.

Proposed method

  • Construction of MENYO-20k, a parallel corpus of 20,000 sentence pairs across multiple domains (e.g., news, social media, education) in English and Yoruba.
  • Application of standardized train-test splits to ensure reproducibility and fair evaluation across models.
  • Fine-tuning of pre-trained multilingual neural machine translation models (e.g., M2M-100, Google multilingual NMT) on the MENYO-20k training split.
  • Evaluation of translation quality using BLEU scores on the standardized test split.
  • Comparison of fine-tuned models against the original pre-trained models to assess performance gains.
  • Use of domain-specific data to enhance model adaptation and improve zero- and few-shot generalization.

Experimental results

Research questions

  • RQ1Can a multi-domain parallel corpus significantly improve neural machine translation performance for low-resource language pairs like Yoruba–English?
  • RQ2How does fine-tuning generic multilingual models on MENYO-20k compare to using the pre-trained models directly in zero- and few-shot settings?
  • RQ3To what extent does domain diversity in the training data contribute to improved translation quality on low-resource language pairs?
  • RQ4Does MENYO-20k enable more reliable and standardized benchmarking of low-resource MT systems compared to existing datasets?
  • RQ5What performance gains are achievable through domain-adaptive fine-tuning on MENYO-20k for the en2yo translation direction?

Key findings

  • Fine-tuning generic multilingual models on MENYO-20k yields a BLEU score improvement of +9.9 over Facebook's M2M-100 for English-to-Yoruba translation.
  • A BLEU gain of +8.6 is achieved when fine-tuning on MENYO-20k compared to Google's multilingual NMT model in the en2yo direction.
  • The proposed benchmarks on MENYO-20k consistently outperform pre-trained multilingual models across all evaluation settings.
  • The multi-domain nature of MENYO-20k contributes to better generalization and robustness in low-resource translation scenarios.
  • Standardized train-test splits in MENYO-20k enable reliable and reproducible evaluation of low-resource MT systems.
  • The results demonstrate that domain-adaptive fine-tuning is highly effective for improving translation quality in low-resource settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.