Skip to main content
QUICK REVIEW

[Paper Review] Evaluation of Text Generation: A Survey

Aslı Çelikyılmaz, Elizabeth Clark|arXiv (Cornell University)|Jun 26, 2020
Topic Modeling295 references194 citations
TL;DR

The paper surveys evaluation methods for natural language generation, grouping them into human-centric, automatic (untrained), and machine-learned metrics, and discusses challenges, tasks, and future directions with example evaluations in summarization and long text generation.

ABSTRACT

The paper surveys evaluation methods of natural language generation (NLG) systems that have been developed in the last few years. We group NLG evaluation methods into three categories: (1) human-centric evaluation metrics, (2) automatic metrics that require no training, and (3) machine-learned metrics. For each category, we discuss the progress that has been made and the challenges still being faced, with a focus on the evaluation of recently proposed NLG tasks and neural NLG models. We then present two examples for task-specific NLG evaluations for automatic text summarization and long text generation, and conclude the paper by proposing future research directions.

Motivation & Objective

  • Motivate the need for robust evaluation of NLG, especially for neural generation systems.
  • Categorize evaluation methods into three families and analyze their progress and challenges.
  • Discuss task-specific evaluation examples (automatic summarization and long-text generation).
  • Propose directions for future research in NLG evaluation to improve comparability and reliability.

Proposed method

  • Classify evaluation methods into three categories: human-centric, untrained automatic metrics, and machine-learned metrics.
  • Review the strengths and limitations of each category in the context of neural NLG systems.
  • Highlight common evaluation dimensions like fluency, adequacy, factuality, and coherence and how they are measured.
  • Illustrate evaluation applications through task-specific examples in automatic summarization and long-text generation.

Experimental results

Research questions

  • RQ1What are the main evaluation paradigms for NLG, and how do they compare in terms of reliability, cost, and scalability?
  • RQ2What progress has been made in human-centric, automatic, and machine-learned evaluation metrics for neural NLG systems?
  • RQ3What are the challenges and future directions for evaluating recent NLG tasks and models?

Key findings

  • Human-centric evaluation remains the gold standard but is expensive and inconsistent across studies.
  • Untrained automatic metrics are prevalent and rely on surface similarities such as n-grams and distributional similarity, but may not align well with human judgments.
  • Machine-learned metrics can model human judgments but require training data and careful calibration to avoid biases.
  • The paper provides examples for task-specific evaluation in automatic summarization and long-text generation to illustrate practical applications and gaps in current metrics.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.