Skip to main content
QUICK REVIEW

[论文解读] Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation

Yixin Liu, Alexander R. Fabbri|arXiv (Cornell University)|Dec 15, 2022
Topic Modeling被引用 4
一句话总结

本文提出了鲁棒摘要评估(RoSE)基准,这是一个基于原子内容单元(ACUs)的大规模人工评估数据集,实现了高一致性的标注者间一致性。研究结果表明,基于ACU的评估方法能产生更稳定的统计结果,揭示了如GPTScore等大语言模型(LLM)指标会过度拟合于与输入无关的标注者偏好,并为跨多种摘要系统评估自动指标提供了可靠基础。

ABSTRACT

Human evaluation is the foundation upon which the evaluation of both summarization systems and automatic metrics rests. However, existing human evaluation studies for summarization either exhibit a low inter-annotator agreement or have insufficient scale, and an in-depth analysis of human evaluation is lacking. Therefore, we address the shortcomings of existing summarization evaluation along the following axes: (1) We propose a modified summarization salience protocol, Atomic Content Units (ACUs), which is based on fine-grained semantic units and allows for a high inter-annotator agreement. (2) We curate the Robust Summarization Evaluation (RoSE) benchmark, a large human evaluation dataset consisting of 22,000 summary-level annotations over 28 top-performing systems on three datasets. (3) We conduct a comparative study of four human evaluation protocols, underscoring potential confounding factors in evaluation setups. (4) We evaluate 50 automatic metrics and their variants using the collected human annotations across evaluation protocols and demonstrate how our benchmark leads to more statistically stable and significant results. The metrics we benchmarked include recent methods based on large language models (LLMs), GPTScore and G-Eval. Furthermore, our findings have important implications for evaluating LLMs, as we show that LLMs adjusted by human feedback (e.g., GPT-3.5) may overfit unconstrained human evaluation, which is affected by the annotators' prior, input-agnostic preferences, calling for more robust, targeted evaluation methods.

研究动机与目标

  • 解决现有摘要人工评估研究中互评一致性低和规模不足的问题。
  • 开发一种基于语义内容单元的更客观、细粒度的人工评估协议。
  • 创建一个大规模、鲁棒的人工评估基准(RoSE),用于评估多种数据集上的最先进摘要系统。
  • 研究不同人工评估协议对自动指标性能和可靠性的影响。
  • 在控制良好、统计功效高的条件下,评估近期基于大语言模型的指标(如GPTScore、G-Eval)的稳健性。

提出的方法

  • 提出原子内容单元(ACU)协议,将摘要分解为细粒度的语义单元,以提高标注一致性和标注者间的一致性。
  • 构建了RoSE基准,涵盖CNN/DailyMail、XSum和SamSum数据集上28个顶尖系统共22,000个摘要级别的标注。
  • 结合内部团队与众包标注,确保人工评估的可靠性和可扩展性。
  • 对比四种人工评估协议:无参考、有参考、基于ACU和LitePyramid,以评估协议对结果的影响。
  • 通过大样本量实现高统计功效,从而获得更紧密的置信区间和更显著的指标比较结果。
  • 在不同协议下评估50种自动指标,包括基于大语言模型的方法(如GPTScore、G-Eval),以评估其与人工判断的相关性。

实验结果

研究问题

  • RQ1基于细粒度内容单元的人工评估协议是否能实现比现有方法更高的标注者间一致性?
  • RQ2不同的人工评估协议(如无参考与有参考)如何影响摘要系统和指标的性能感知?
  • RQ3基于大语言模型的指标(如GPTScore)在多大程度上会过度拟合于标注者的先验偏好和与输入无关的偏好?
  • RQ4高统计功效、大规模的人工评估基准是否能带来更稳定、更可靠的自动指标评估结果?
  • RQ5传统自动指标与基于大语言模型的自动指标在不同评估协议下,与稳健人工评估的相关性如何?

主要发现

  • ACU协议实现了高标注者间一致性,使系统评估更加稳定且可复现。
  • 无参考人工评估与标注者对输入无关的偏好高度相关,尤其倾向于更长的摘要。
  • 零样本大语言模型(如GPT-3)在无参考评估中表现优于微调模型,但在有参考评估中表现更差。
  • 基于大语言模型的指标(如GPTScore和G-Eval)显著过度拟合于标注者偏好,在采用稳健人工标准评估时表现不佳。
  • RoSE基准支持更强大的统计评估,使置信区间更紧密,指标比较研究的显著性更高。
  • 传统指标(如BERTScore和ROUGE)与基于ACU的人工评估在系统层面的相关性,强于较新的基于大语言模型的指标。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。