[Paper Review] Perspectives on Large Language Models for Relevance Judgment
This perspectives paper debates using large language models (LLMs) for relevance judgments in IR, presents a human–machine collaboration spectrum, and reports a preliminary pilot comparing LLM judgments with human assessors. It discusses open issues, risks, and potential paths toward fully or partially automated test collections.
When asked, large language models (LLMs) like ChatGPT claim that they can assist with relevance judgments but it is not clear whether automated judgments can reliably be used in evaluations of retrieval systems. In this perspectives paper, we discuss possible ways for LLMs to support relevance judgments along with concerns and issues that arise. We devise a human--machine collaboration spectrum that allows to categorize different relevance judgment strategies, based on how much humans rely on machines. For the extreme point of "fully automated judgments", we further include a pilot experiment on whether LLM-based relevance judgments correlate with judgments from trained human assessors. We conclude the paper by providing opposing perspectives for and against the use of~LLMs for automatic relevance judgments, and a compromise perspective, informed by our analyses of the literature, our preliminary experimental evidence, and our experience as IR researchers.
Motivation & Objective
- Motivate and frame the evaluation challenges in IR per the Cranfield paradigm and the cost of human judgments.
- Propose a spectrum of human–machine collaboration for relevance judgments to assess feasibility and costs.
- Survey existing approaches (manual, crowd, AI-assisted, fully automated) and their trade-offs.
- Provide preliminary empirical evidence on agreement between LLMs and human judgments.
- Outline open issues, risks, and potential future directions for LLM-based relevance assessment.
Proposed method
- Review and synthesize literature on relevance judgments and automatic assistance.
- Propose a four-level human–machine collaboration spectrum from manual to fully automated judgments.
- Conduct a pilot feasibility experiment comparing LLM-based judgments (GPT-3.5 and YouChat) against human assessors on TREC-8 and TREC-DL 2021.
- Re-judge TREC-DL 2021 using GPT-3.5 with a few-shot prompt setup and compare with original human judgments.
- Discuss biases, factuality, and reliability concerns of LLM-based judgments and human verification strategies.
Experimental results
Research questions
- RQ1Can LLMs produce relevance judgments that meaningfully align with trained human assessors across different test collections?
- RQ2What is the cost-quality trade-off of using LLMs for relevance judgments compared to human assessors?
- RQ3How should human–machine collaboration be structured to maximize reliability and efficiency in relevance judgments?
- RQ4What open risks (bias, hallucination, truthfulness) arise when relying on LLMs for test collections?
- RQ5Is a fully automated LLM-based evaluation feasible, and under what conditions?
Key findings
- LLMs show partial agreement with human assessors, with higher alignment on certain non-relevant cases and more mixed results on relevant cases depending on the collection and model.
- GPT-3.5 achieved 0.38 Cohen’s kappa for relevant vs non-relevant on TREC-8 in one setup, while YouChat showed lower agreement in the same task.
- On TREC-DL 2021, YouChat achieved higher agreement for highly relevant (grade 3) cases (0.49 kappa) than for non-relevant cases (0.42 agreement in binarized form), indicating variable performance across relevance grades.
- A more favorable alignment was observed for YouChat on highly relevant question–passage pairs (96 of 100) compared to non-relevant ones (42 of 100) in TREC-DL 2021.
- The authors demonstrate a cost difference in the re-judging experiment for TREC-DL 2021, noting GPT-3.5 judgments cost about USD 0.01 per judgment and total expenditure of USD 111.90 in their setup.
- The paper highlights multiple open issues including bias, factuality, reasoning, and the need for quality assurance in LLM-based judgments, as well as the potential for personalized or diversified LLMs to reduce correlation across models.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.