[论文解读] Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
本论文评估将LLMs作为领域特定任务的判决者在膳食学和精神健康领域,发现与主题专家(SME)的一致性有限,并倡导持续的 SME 参与。
The potential of using Large Language Models (LLMs) themselves to evaluate LLM outputs offers a promising method for assessing model performance across various contexts. Previous research indicates that LLM-as-a-judge exhibits a strong correlation with human judges in the context of general instruction following. However, for instructions that require specialized knowledge, the validity of using LLMs as judges remains uncertain. In our study, we applied a mixed-methods approach, conducting pairwise comparisons in which both subject matter experts (SMEs) and LLMs evaluated outputs from domain-specific tasks. We focused on two distinct fields: dietetics, with registered dietitian experts, and mental health, with clinical psychologist experts. Our results showed that SMEs agreed with LLM judges 68% of the time in the dietetics domain and 64% in mental health when evaluating overall preference. Additionally, the results indicated variations in SME-LLM agreement across domain-specific aspect questions. Our findings emphasize the importance of keeping human experts in the evaluation process, as LLMs alone may not provide the depth of understanding required for complex, knowledge specific tasks. We also explore the implications of LLM evaluations across different domains and discuss how these insights can inform the design of evaluation workflows that ensure better alignment between human experts and LLMs in interactive systems.
研究动机与目标
- 评估基于LLM的评估在领域特定、专业知识任务中与主题专家(SME)判断的一致性。
- 调查驱动LLM评判与SME之间在膳食学和精神健康领域的一致性/不一致性的因素。
- 考察专家角色对LLM评判与SME对齐的影响。
- 分析SME与LLM在评估差异方面的定性解释。
提出的方法
- 为膳食学和精神健康 curated 25 条领域特定指令。
- 对每条指令使用SME和LLM判决进行成对比较评估,比较两个模型输出。
- 使用专家角色提示测试对一致性的影响。
- 应用 AlpacaEval 框架进行带解释的LLM排序。
- 对排序解释进行自反性主题分析以识别主题。

实验结果
研究问题
- RQ1RQ1: LLM作为判决者的评估与领域特定任务的SME评估相比如何?
- RQ2RQ2: 影响LLMs与SMEs评估差异及解释的因素有哪些?
主要发现
- SMEs在膳食学中与LLM判决者的一致性为68%,在精神健康中为64%,总体偏好。
- SMEs彼此之间的一致性为精神健康72%(精神健康)和膳食学75%(膳食学)。
- 专家角色提示将SME–LLM的一致性提高了约4%(一般偏好)。
- 在领域特定方面的一致性存在差异,且在多个类别中精神健康的对齐通常高于膳食学。
- SMEs强调准确性、最新证据、专业标准和清晰的沟通;LLMs常强调遵循指令和细节。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。