Skip to main content
QUICK REVIEW

[论文解读] From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge

Dawei Li, Bocheng Jiang|arXiv (Cornell University)|Nov 25, 2024
Legal Education and Practice Innovations被引用 10
一句话总结

This paper surveys the LLM-as-a-judge paradigm, defining input/output formats, presenting a three-dimension taxonomy (what/how/where to judge), compiling evaluation benchmarks, and outlining key challenges and future directions.

ABSTRACT

Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or small model-based, often fall short in open-ended and dynamic scenarios. Recent advancements in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm, where LLMs are leveraged to perform scoring, ranking, or selection for various machine learning evaluation scenarios. This paper presents a comprehensive survey of LLM-based judgment and assessment, offering an in-depth overview to review this evolving field. We first provide the definition from both input and output perspectives. Then we introduce a systematic taxonomy to explore LLM-as-a-judge along three dimensions: what to judge, how to judge, and how to benchmark. Finally, we also highlight key challenges and promising future directions for this emerging area. More resources on LLM-as-a-judge are on the website: https://llm-as-a-judge.github.io and https://github.com/llm-as-a-judge/Awesome-LLM-as-a-judge.

研究动机与目标

  • Formalize LLM-as-a-judge definitions from input and output perspectives.
  • Propose a comprehensive taxonomy for judging attributes, methodologies, and applications.
  • Compile and summarize benchmarks for evaluating LLM-based judgments.
  • Identify challenges and highlight promising directions for future research.

提出的方法

  • Define input formats (point-wise, pair/list-wise) and output formats (score, ranking, selection).
  • Develop a three-dimensional taxonomy: what to judge (attributes), how to judge (tuning and prompting), where to judge (applications).
  • Survey attributes like helpfulness, harmlessness, reliability, relevance, feasibility, and overall quality.
  • Summarize tuning techniques (data sources, supervised fine-tuning, preference learning) and prompting strategies (swapping, rule augmentation, multi-agent collaboration).
  • List and categorize existing benchmarks for evaluating LLM-based judgment across tasks.
Figure 1: Overview of various input and output formats of LLM-as-a-judge.
Figure 1: Overview of various input and output formats of LLM-as-a-judge.

实验结果

研究问题

  • RQ1What attributes can LLMs judge effectively and how are these attributes defined and measured?
  • RQ2What tuning and prompting methodologies enable robust LLM-based judging across tasks?
  • RQ3In which applications are LLM-as-a-judge approaches currently employed and how are they benchmarked?
  • RQ4What are the main challenges and open directions for LLM-as-a-judge research?

主要发现

  • The paper presents a detailed taxonomy addressing what, how, and where to judge with LLMs.
  • It catalogs a wide range of attributes (e.g., helpfulness, harmlessness, reliability, relevance, feasibility, and overall quality).
  • It reviews tuning methods (SFT, preference learning, synthetic data) and prompting tricks (swapping, rule augmentation, multi-agent setups).
  • It compiles benchmarks and maps LLM-as-a-judge usage to evaluation, alignment, retrieval, and reasoning applications.
  • It discusses challenges such as bias, vulnerability, dynamics of judgment, and human–LLM co-judgment.
Figure 3: LLMs are capable of judging various attributes.
Figure 3: LLMs are capable of judging various attributes.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。