Skip to main content
QUICK REVIEW

[论文解读] A Survey on Evaluation of Large Language Models

Yupeng Chang, Xu Wang|arXiv (Cornell University)|Jul 6, 2023
Topic Modeling被引用 195
一句话总结

本论文综述大语言模型(LLMs)的评估方法,涵盖评估什么、在哪里评估、以及如何评估,突出任务、基准和挑战。

ABSTRACT

Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate, where to evaluate, and how to evaluate. Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, educations, natural and social sciences, agent applications, and other areas. Secondly, we answer the `where' and `how' questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey.

研究动机与目标

  • 总结现有的对LLMs在NLP、推理、伦理、教育、科学与应用领域的评估任务。
  • 分析用于评估LLM性能的评估数据集与基准。
  • 讨论评估方法学,包括自动评估与人工评估,并指出优点与局限性。
  • 突出评估LLMs的重大挑战与未来方向,追求原理性、稳健性与全面性。

提出的方法

  • 将LLM评估划分为三个维度:评估什么、在哪里评估、以及如何评估。
  • 回顾NLP任务(NLU、NLG、推理、多语种、事实性)及其他领域(医学、伦理学、社会科学、代理应用)。
  • 整理用于LLM评估的一般与特定基准与数据集(GLUE、MMLU、BIG-bench 等)。
  • 考察自动评估与人工评估的方法及它们在LLM评估中的作用。
  • 讨论LLM评估的重大挑战与开源资源(开放仓库)。
Figure 3. The evaluation process of AI models.
Figure 3. The evaluation process of AI models.

实验结果

研究问题

  • RQ1用于评估LLMs的任务有哪些?它们对LLMs的优缺点揭示了什么?
  • RQ2LLMs在哪些数据集与基准上被评估?哪些基准能较好地反映它们的能力?
  • RQ3LLMs的评估如何进行(自动 vs. 人工、协议设计),当前评估方法的局限性有哪些?
  • RQ4评估LLMs的重大挑战和未来方向是什么?
  • RQ5从中可以得出哪些见解以指导更强健、更可信的LLMs的开发?

主要发现

  • LLMs在许多NLP任务上表现强劲,但在某些推理和语义理解方面仍存在弱点。
  • 评估基准在范围上各不相同,可能无法充分捕捉新兴能力或LLMs的安全性考量。
  • 自动评估和人工评估都很重要,但各自存在限制,影响可靠性与解读。
  • 需要统一、原理性评估框架,覆盖通用与领域特定任务、鲁棒性与可信度。
  • 开源材料和持续的基准开发对LLM评估的协同进展至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。