Skip to main content
QUICK REVIEW

[Paper Review] LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Haitao Li, Qian Dong|arXiv (Cornell University)|Dec 7, 2024
Legal Education and Practice Innovations21 citations
TL;DR

This paper surveys the paradigm of using Large Language Models as evaluators (LLMs-as-judges) across functionality, methodology, applications, meta-evaluation, and limitations, and provides an open-source resource for the community.

ABSTRACT

The rapid advancement of Large Language Models (LLMs) has driven their expanding application across various fields. One of the most promising applications is their role as evaluators based on natural language responses, referred to as ''LLMs-as-judges''. This framework has attracted growing attention from both academia and industry due to their excellent effectiveness, ability to generalize across tasks, and interpretability in the form of natural language. This paper presents a comprehensive survey of the LLMs-as-judges paradigm from five key perspectives: Functionality, Methodology, Applications, Meta-evaluation, and Limitations. We begin by providing a systematic definition of LLMs-as-Judges and introduce their functionality (Why use LLM judges?). Then we address methodology to construct an evaluation system with LLMs (How to use LLM judges?). Additionally, we investigate the potential domains for their application (Where to use LLM judges?) and discuss methods for evaluating them in various contexts (How to evaluate LLM judges?). Finally, we provide a detailed analysis of the limitations of LLM judges and discuss potential future directions. Through a structured and comprehensive analysis, we aim aims to provide insights on the development and application of LLMs-as-judges in both research and practice. We will continue to maintain the relevant resource list at https://github.com/CSHaitao/Awesome-LLMs-as-Judges.

Motivation & Objective

  • Define and formalize the LLMs-as-judges paradigm and its evaluation framework.
  • Systematically analyze current research from five perspectives: Functionality, Methodology, Applications, Meta-evaluation, and Limitations.
  • Identify challenges, opportunities, and future directions to guide research and practice.
  • Provide an open-source repository to foster community collaboration and best practices.

Proposed method

  • Classify evaluation setups into Single-LLM, Multi-LLM, and Hybrid (human-AI) configurations.
  • Describe inputs (Evaluation Type, Criteria, References) and outputs (Evaluation Result, Explanation, Feedback) of LLM judges.
  • Detail evaluation modes (pointwise, pairwise, listwise) and how criteria and references influence judgments.
  • Survey methodological approaches including prompting, tuning, data construction, and multi-LLM aggregation.
  • Discuss meta-evaluation benchmarks and metrics used to assess LLM-based evaluation performance.

Experimental results

Research questions

  • RQ1What are the core components and definitions of LLMs-as-judges?
  • RQ2How are LLM-based evaluators constructed and configured across single, multi-LLM, and human–AI hybrid setups?
  • RQ3What domains, tasks, and criteria are most impacted by LLM-based evaluation methods?
  • RQ4How should LLM judges be evaluated themselves (meta-evaluation) and what are their limitations?
  • RQ5What future directions can improve efficiency, effectiveness, reliability, and fairness of LLM-based evaluation?

Key findings

  • LLMs-as-judges offer flexible evaluation criteria and interpretable feedback that can scale and generalize across tasks.
  • Evaluation outputs typically include a main result plus explanations and actionable feedback, enabling transparent assessments.
  • There are notable biases, prompt-dependence, and transitivity issues that affect reliability and fairness of LLM judgments.
  • The survey maps methods from prompt-based to tuning-based and multi-LLM aggregation, highlighting trade-offs in cost and robustness.
  • An open-source resource (Awesome-LLMs-as-Judges) is provided to support ongoing collaboration and standardization.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.