Skip to main content
QUICK REVIEW

[Paper Review] Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks

Annalisa Szymanski, Noah Ziems|arXiv (Cornell University)|Oct 26, 2024
Quality and Management Systems5 citations
TL;DR

The paper evaluates LLMs as judges for domain-specific tasks in dietetics and mental health, finding limited SME-LLM agreement and advocating ongoing SME involvement.

ABSTRACT

The potential of using Large Language Models (LLMs) themselves to evaluate LLM outputs offers a promising method for assessing model performance across various contexts. Previous research indicates that LLM-as-a-judge exhibits a strong correlation with human judges in the context of general instruction following. However, for instructions that require specialized knowledge, the validity of using LLMs as judges remains uncertain. In our study, we applied a mixed-methods approach, conducting pairwise comparisons in which both subject matter experts (SMEs) and LLMs evaluated outputs from domain-specific tasks. We focused on two distinct fields: dietetics, with registered dietitian experts, and mental health, with clinical psychologist experts. Our results showed that SMEs agreed with LLM judges 68% of the time in the dietetics domain and 64% in mental health when evaluating overall preference. Additionally, the results indicated variations in SME-LLM agreement across domain-specific aspect questions. Our findings emphasize the importance of keeping human experts in the evaluation process, as LLMs alone may not provide the depth of understanding required for complex, knowledge specific tasks. We also explore the implications of LLM evaluations across different domains and discuss how these insights can inform the design of evaluation workflows that ensure better alignment between human experts and LLMs in interactive systems.

Motivation & Objective

  • Assess how LLM-based evaluations align with subject matter expert (SME) judgments in domain-specific, expert-knowledge tasks.
  • Investigate factors driving agreement/disagreement between SMEs and LLM judges across dietetics and mental health domains.
  • Examine how expert personas influence alignment between LLM judges and SMEs.
  • Analyze qualitative explanations from both SMEs and LLMs to understand evaluation differences.

Proposed method

  • Curated 25 domain-specific instructions for dietetics and mental health.
  • Compare two model outputs per instruction using pairwise evaluation by SMEs and an LLM judge.
  • Use expert persona prompts to test impact on agreement.
  • Apply AlpacaEval framework for LLM-based ranking with explanations.
  • Perform reflexive thematic analysis on ranking explanations to identify themes.
Figure 1. Comparison of Agreement between LLM Judge and SMEs versus Lay Users for Dietetics and Mental Health. The lay users show significantly more agreement with the LLM Judge using a standard persona than the LLM Judge using an expert persona. Conversely, the SMEs show more agreement with the LLM
Figure 1. Comparison of Agreement between LLM Judge and SMEs versus Lay Users for Dietetics and Mental Health. The lay users show significantly more agreement with the LLM Judge using a standard persona than the LLM Judge using an expert persona. Conversely, the SMEs show more agreement with the LLM

Experimental results

Research questions

  • RQ1RQ1: How does LLM-as-a-Judge evaluation compare to SME evaluations for domain-specific tasks?
  • RQ2RQ2: What factors contribute to evaluation differences and explanations between LLMs and SMEs?

Key findings

  • SMEs agreed with LLM judges 68% of the time in dietetics and 64% in mental health for overall preference.
  • SMEs agreed with each other 72% (mental health) and 75% (dietetics).
  • Expert persona prompts improved SME–LLM agreement by about 4% for general preferences.
  • Agreement varied across domain-specific aspects and was generally higher for mental health than dietetics in several categories.
  • SMEs prioritized accuracy, up-to-date evidence, professional standards, and clear communication; LLMs often emphasized instruction-following and detail.
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.