Skip to main content
QUICK REVIEW

[论文解读] Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Jiaan Wang, Yunlong Liang|arXiv (Cornell University)|Mar 7, 2023
Topic Modeling被引用 12
一句话总结

论文初步评估 ChatGPT 作为通用 NLG 评估器,在多个元评估数据集上,与人工判断呈现出高相关性,且受提示与数据集偏差影响表现。

ABSTRACT

Recently, the emergence of ChatGPT has attracted wide attention from the computational linguistics community. Many prior studies have shown that ChatGPT achieves remarkable performance on various NLP tasks in terms of automatic evaluation metrics. However, the ability of ChatGPT to serve as an evaluation metric is still underexplored. Considering assessing the quality of natural language generation (NLG) models is an arduous task and NLG metrics notoriously show their poor correlation with human judgments, we wonder whether ChatGPT is a good NLG evaluation metric. In this report, we provide a preliminary meta-evaluation on ChatGPT to show its reliability as an NLG metric. In detail, we regard ChatGPT as a human evaluator and give task-specific (e.g., summarization) and aspect-specific (e.g., relevance) instruction to prompt ChatGPT to evaluate the generated results of NLG models. We conduct experiments on five NLG meta-evaluation datasets (including summarization, story generation and data-to-text tasks). Experimental results show that compared with previous automatic metrics, ChatGPT achieves state-of-the-art or competitive correlation with human judgments in most cases. In addition, we find that the effectiveness of the ChatGPT evaluator might be influenced by the creation method of the meta-evaluation datasets. For the meta-evaluation datasets which are created greatly depending on the reference and thus are biased, the ChatGPT evaluator might lose its effectiveness. We hope our preliminary study could prompt the emergence of a general-purposed reliable NLG metric.

研究动机与目标

  • 激励并评估 ChatGPT 是否能作为通用 NLG 评估指标。
  • 指出传统自动度量的局限性以及作为无参考与有参考评估者的潜力。
  • 研究任务特定和方面特定的提示如何影响 ChatGPT 对各任务(摘要、故事生成、数据转文本)的 NLG 输出评判。
  • 检验数据集构造偏差如何影响 ChatGPT 作为评估指标的有效性。

提出的方法

  • 将 ChatGPT 视为人类评估者,应用任务特定和方面特定的提示以产生分数或评分。
  • 将基于 ChatGPT 的评估与标准自动度量(ROUGE、BERTScore、MoverScore、PRISM、BARTScore 等)进行比较。
  • 使用无参考提示(DA 和星级评分)以及有参考提示(带黄金参考)来引导评分。
  • 在五个元评估数据集上进行评估,涵盖摘要、故事生成和数据转文本任务。
  • 使用样本层面和数据集层面的指标(如 Spearman、Pearson、Kendall)分析与人工判断的相关性。
  • 评估提示设计和数据集构造偏差对 ChatGPT 评估器性能的影响。

实验结果

研究问题

  • RQ1ChatGPT 是否能够作为跨越多任务的通用 NLG 评估器,与人工判断保持相关性?
  • RQ2任务特定和方面特定的提示如何影响 ChatGPT 的评估可靠性?
  • RQ3元评估数据集中的偏差是否影响 ChatGPT 作为 NLG 指标的有效性?
  • RQ4在摘要、故事生成和数据转文本任务中,ChatGPT 与既有自动度量相比如何?

主要发现

  • ChatGPT 在多个元评估数据集上与人工判断呈现出高相关性,尤其在故事生成和摘要情境中。
  • 在多项任务中,ChatGPT 与人工判断的相关性往往优于传统自动度量,显示出作为通用 NLG 指标的潜力。
  • 评估者的有效性对提示设计敏感,不同任务和方面需要定制化提示。
  • 偏向参考相似性的数据集可能降低 ChatGPT 的有效性,因为词汇偏见可能偏向参考信号。
  • 在数据转文本评估中,ChatGPT 展现出竞争力,显示出在摘要和故事生成之外的广泛适用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。