[论文解读] LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
本论文综述将大型语言模型作为评估者(LLMs-as-judges)的范式,在功能性、方法论、应用、元评估与局限性方面,并为社区提供一个开源资源。
The rapid advancement of Large Language Models (LLMs) has driven their expanding application across various fields. One of the most promising applications is their role as evaluators based on natural language responses, referred to as ''LLMs-as-judges''. This framework has attracted growing attention from both academia and industry due to their excellent effectiveness, ability to generalize across tasks, and interpretability in the form of natural language. This paper presents a comprehensive survey of the LLMs-as-judges paradigm from five key perspectives: Functionality, Methodology, Applications, Meta-evaluation, and Limitations. We begin by providing a systematic definition of LLMs-as-Judges and introduce their functionality (Why use LLM judges?). Then we address methodology to construct an evaluation system with LLMs (How to use LLM judges?). Additionally, we investigate the potential domains for their application (Where to use LLM judges?) and discuss methods for evaluating them in various contexts (How to evaluate LLM judges?). Finally, we provide a detailed analysis of the limitations of LLM judges and discuss potential future directions. Through a structured and comprehensive analysis, we aim aims to provide insights on the development and application of LLMs-as-judges in both research and practice. We will continue to maintain the relevant resource list at https://github.com/CSHaitao/Awesome-LLMs-as-Judges.
研究动机与目标
- 定义并形式化LLMs-as-judges范式及其评估框架。
- 从五个角度系统分析当前研究:功能性、方法论、应用、元评估与局限性。
- 识别挑战、机会与未来方向,以指导研究与实践。
- 提供一个开源仓库,促进社区协作与最佳实践。
提出的方法
- 将评估设置分类为单一LLM、多LLM以及混合(人机)配置。
- 描述LLM评判者的输入(评估类型、标准、参考)和输出(评估结果、解释、反馈)。
- 详述评估模式(逐点、逐对、逐列)以及标准和参考如何影响判断。
- 调查方法论方法,包括提示、微调、数据构建以及多LLM聚合。
- 讨论元评估基准和用于评估基于LLM的评估性能的度量。
实验结果
研究问题
- RQ1LLMs-as-judges的核心组成要素与定义是什么?
- RQ2LLM评估人如何在单一、多人LLM、以及人机混合设置中构建与配置?
- RQ3哪些领域、任务和标准最受基于LLM的评估方法的影响?
- RQ4基于LLM的评估者自身应如何进行评估(元评估),它们的局限性是什么?
- RQ5哪些未来方向可以提高基于LLM的评估的效率、有效性、可靠性和公平性?
主要发现
- LLMs-as-judges 提供灵活的评估标准和可解释的反馈,能够跨任务扩展和泛化。
- 评估输出通常包含一个主结果以及解释和可操作的反馈,从而实现透明的评估。
- 存在显著的偏见、提示依赖性和传递性问题,这些会影响LLM判断的可靠性和公平性。
- 该综述将方法从基于提示到微调再到多LLM聚合进行了映射,强调成本与鲁棒性之间的权衡。
- 提供一个开源资源(Awesome-LLMs-as-Judges),以支持持续协作和标准化。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。