Skip to main content
QUICK REVIEW

[Paper Review] The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

Sebastian Gehrmann, Tosin Adewumi|arXiv (Cornell University)|Feb 2, 2021
Topic Modeling111 references52 citations
TL;DR

GEM introduces a living, multilingual benchmark for NLG focused on generation, evaluation, and metrics, with open data cards, challenge sets, and a flexible evaluation framework.

ABSTRACT

We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of automated metrics, datasets, and human evaluation standards. Due to this moving target, new models often still evaluate on divergent anglo-centric corpora with well-established, but flawed, metrics. This disconnect makes it challenging to identify the limitations of current models and opportunities for progress. Addressing this limitation, GEM provides an environment in which models can easily be applied to a wide set of tasks and in which evaluation strategies can be tested. Regular updates to the benchmark will help NLG research become more multilingual and evolve the challenge alongside models. This paper serves as the description of the data for which we are organizing a shared task at our ACL 2021 Workshop and to which we invite the entire NLG community to participate.

Motivation & Objective

  • Provide a living, multilingual benchmark ecosystem for NLG that evolves with models and evaluation standards.
  • Enable comprehensive evaluation combining human and automated metrics beyond single-number scores.
  • Promote responsible data usage with data cards and standardized evaluation protocols.
  • Incorporate diverse high-quality datasets across languages and generation tasks to reduce anglocentric bias.
  • Offer challenge sets to probe model behavior and generalization under targeted conditions.

Proposed method

  • Curate an initial set of 11 NLG datasets spanning summarization, dialog, data-to-text, and simplification across 18 languages.
  • Adopt a three-step dataset selection process (proposal, criteria, voting) to maximize utility under resource constraints.
  • Create NLG-specific data cards documenting dataset characteristics, limitations, and real-world use cases.
  • Develop challenge set types (input perturbations, subset splits, time-shift data) to diagnose model behavior beyond i.i.d. test sets.
  • Outline an experimental setup with baselines (e.g., T5, BART, mT5, mBART) and a framework for expanding automated metrics.
  • Position GEM as a living benchmark that replaces solved tasks with harder ones over time and supports new metrics.

Experimental results

Research questions

  • RQ1How can a living, multilingual benchmark better capture the multifaceted goals of NLG evaluation beyond traditional metrics?
  • RQ2What dataset composition, languages, and task mix maximize robustness and generalizability of NLG models?
  • RQ3How can challenge sets reveal model limitations and bias that standard test sets miss?
  • RQ4What standards for data documentation and human evaluation are needed to ensure reproducibility and responsible use?
  • RQ5How do automated metrics correlate with human judgments across diverse NLG tasks and languages?

Key findings

  • GEM proposes a diverse, multilingual dataset suite with 18 languages and tasks like summarization, dialog, data-to-text, and simplification.
  • Datasets are curated with data cards to document limitations and real-world use cases, supporting responsible research.
  • Challenge sets are designed to probe numerical variation, attribute order, typographical errors, backtranslation, and input structure, among others.
  • A living benchmark structure is described, enabling updates to data, test sets, and metrics as the field evolves.
  • Baseline modeling approaches (e.g., T5, BART, mT5, mBART) are discussed to establish a starting point for evaluation, with plans to expand metrics beyond traditional n-gram overlap (BLEU/ROUGE).
  • The paper emphasizes avoiding leaderboard-driven optimization by focusing on in-depth evaluation across human and automatic metrics.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.