[论文解读] CohEval: Benchmarking Coherence Models
本文提出了CohEval,一个用于在合成句子排序任务和三个下游应用(机器翻译、文本摘要和下一句预测)中评估连贯性模型的基准。研究发现,合成任务表现与真实世界连贯性效果之间存在微弱相关性,凸显了改进评估协议的必要性,并由此催生了一个公开排行榜,以推动连贯性建模研究的发展。
Although coherence modeling has come a long way in developing novel models, their evaluation on downstream applications has largely been neglected. With the advancements made by neural approaches in applications such as machine translation, text summarization and dialogue systems, the need for standard coherence evaluation is now more crucial than ever. In this paper, we propose to benchmark coherence models on a number of synthetic and downstream tasks. In particular, we evaluate well-known traditional and neural coherence models on sentence ordering tasks, and also on three downstream applications including coherence evaluation for machine translation, summarization and next utterance prediction. We also show model produced rankings for pre-trained language model outputs as another use-case. Our results demonstrate a weak correlation between the model performances in the synthetic tasks and the downstream applications, motivating alternate evaluation methods for coherence models. This work has led us to create a leaderboard to foster further research in coherence modeling.
研究动机与目标
- 解决NLP应用中连贯性模型缺乏标准化评估的问题。
- 探究在合成连贯性任务上的表现是否能泛化到真实世界的下游应用。
- 开发一个全面的基准,用于在多样化NLP任务中评估连贯性模型。
- 识别当前评估实践中的不足,并促进更可靠的模型评估。
提出的方法
- 提出一个结合合成句子排序任务和真实世界下游应用的基准框架。
- 在合成任务和应用特定任务上评估传统与神经网络连贯性模型。
- 使用预训练语言模型输出来评估模型生成的排序结果,作为额外的应用场景。
- 在多个评估设置下收集并分析模型表现,以比较合成任务与下游任务的结果。
- 创建一个公开排行榜,以追踪并激励连贯性建模领域的进展。
实验结果
研究问题
- RQ1在合成句子排序任务上训练的连贯性模型在真实世界NLP应用中的泛化能力如何?
- RQ2在合成连贯性任务上的表现与在下游连贯性评估任务上的表现之间存在何种相关性?
- RQ3模型对预训练语言模型输出生成的排序结果能否作为连贯性的有效评估信号?
- RQ4当前的评估实践是否充分反映了生成文本在实际场景中的真实连贯性质量?
主要发现
- 模型在合成句子排序任务上的表现与在下游连贯性应用中的表现之间存在微弱相关性。
- 在合成任务上表现良好的模型,往往在机器翻译和摘要等真实世界任务中表现欠佳。
- 这种差异表明,仅靠合成基准不足以评估连贯性模型。
- 所提出的基准揭示了当前评估方法论中的显著差距,迫切需要替代性评估策略。
- 基于此基准创建的公开排行榜有望推动连贯性建模领域的创新与标准化。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。