[Paper Review] Is ChatGPT a Good NLG Evaluator? A Preliminary Study
The paper preliminarily evaluates ChatGPT as a general NLG evaluator, finding high correlations with human judgments on several meta-evaluation datasets, with performance influenced by prompts and dataset biases.
Recently, the emergence of ChatGPT has attracted wide attention from the computational linguistics community. Many prior studies have shown that ChatGPT achieves remarkable performance on various NLP tasks in terms of automatic evaluation metrics. However, the ability of ChatGPT to serve as an evaluation metric is still underexplored. Considering assessing the quality of natural language generation (NLG) models is an arduous task and NLG metrics notoriously show their poor correlation with human judgments, we wonder whether ChatGPT is a good NLG evaluation metric. In this report, we provide a preliminary meta-evaluation on ChatGPT to show its reliability as an NLG metric. In detail, we regard ChatGPT as a human evaluator and give task-specific (e.g., summarization) and aspect-specific (e.g., relevance) instruction to prompt ChatGPT to evaluate the generated results of NLG models. We conduct experiments on five NLG meta-evaluation datasets (including summarization, story generation and data-to-text tasks). Experimental results show that compared with previous automatic metrics, ChatGPT achieves state-of-the-art or competitive correlation with human judgments in most cases. In addition, we find that the effectiveness of the ChatGPT evaluator might be influenced by the creation method of the meta-evaluation datasets. For the meta-evaluation datasets which are created greatly depending on the reference and thus are biased, the ChatGPT evaluator might lose its effectiveness. We hope our preliminary study could prompt the emergence of a general-purposed reliable NLG metric.
Motivation & Objective
- Motivate and assess whether ChatGPT can serve as a general NLG evaluation metric.
- Call out the limitations of traditional automatic metrics and potential of ChatGPT as a reference-free and reference-based evaluator.
- Investigate how task-specific and aspect-specific prompts influence ChatGPT's judging of NLG outputs across tasks (summarization, story generation, data-to-text).
- Examine how dataset construction biases impact the effectiveness of ChatGPT as an evaluation metric.
Proposed method
- Treat ChatGPT as a human evaluator and apply task-specific and aspect-specific prompts to produce scores or ratings.
- Compare ChatGPT-based evaluations against standard automatic metrics (ROUGE, BERTScore, MoverScore, PRISM, BARTScore, etc.).
- Use both reference-free prompts (DA and star ratings) and reference-based prompts (with golden references) to guide scoring.
- Evaluate on five meta-evaluation datasets spanning summarization, story generation, and data-to-text tasks.
- Analyze correlations with human judgments using sample-level and dataset-level metrics (e.g., Spearman, Pearson, Kendall).
- Assess the impact of prompt design and dataset construction biases on ChatGPT evaluator performance.
Experimental results
Research questions
- RQ1Is ChatGPT capable of correlating with human judgments as a general NLG evaluator across multiple tasks?
- RQ2How do task-specific and aspect-specific prompts affect ChatGPT's evaluation reliability?
- RQ3Do biases in meta-evaluation datasets influence the effectiveness of ChatGPT as an NLG metric?
- RQ4How does ChatGPT compare to established automatic metrics across summarization, story generation, and data-to-text tasks?
Key findings
- ChatGPT achieves high correlations with human judgments on several meta-evaluation datasets, notably in story generation and summarization contexts.
- ChatGPT often outperforms traditional automatic metrics in correlation with human judgments across multiple tasks, indicating potential as a general NLG metric.
- The evaluator’s effectiveness is sensitive to prompt design, with different tasks and aspects requiring tailored prompts.
- Datasets biased toward reference similarity can diminish ChatGPT’s effectiveness, as lexical biases can favor reference-based signals.
- ChatGPT shows competitive performance in data-to-text evaluation, indicating broad applicability beyond summarization and story generation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.