[Paper Review] A Survey on Evaluation of Large Language Models
This paper surveys evaluation methods for large language models (LLMs) across what to evaluate, where to evaluate, and how to evaluate, highlighting tasks, benchmarks, and challenges.
Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate, where to evaluate, and how to evaluate. Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, educations, natural and social sciences, agent applications, and other areas. Secondly, we answer the `where' and `how' questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey.
Motivation & Objective
- Summarize existing evaluation tasks for LLMs across NLP, reasoning, ethics, education, science, and applications.
- Analyze evaluation datasets and benchmarks used to assess LLM performance.
- Discuss evaluation methodologies, including automatic and human evaluations, and identify strengths and limitations.
- Highlight grand challenges and future directions for principled, robust, and comprehensive LLM evaluation.
Proposed method
- Classifies LLM evaluation into three dimensions: what to evaluate, where to evaluate, and how to evaluate.
- Reviews NLP tasks (NLU, NLG, reasoning, multilingual, factuality) and other domains (medical, ethics, social sciences, agent applications).
- Catalogues general and specific benchmarks and datasets used for LLM evaluation (GLUE, MMLU, BIG-bench, etc.).
- Examines automatic vs. human evaluation approaches and their role in LLM assessment.
- Discusses grand challenges and open-source resources for LLM evaluation (open repository).

Experimental results
Research questions
- RQ1What evaluation tasks are used to assess LLMs and what do they reveal about strengths and weaknesses?
- RQ2Where (which datasets and benchmarks) are LLMs evaluated, and what benchmarks capture their capabilities well?
- RQ3How are LLMs evaluated (automatic vs. human, protocol design), and what are the limitations of current evaluation methods?
- RQ4What are the grand challenges and future directions in evaluating LLMs?
- RQ5What insights can be drawn to guide the development of more robust and trustworthy LLMs?
Key findings
- LLMs show strong performance in many NLP tasks but exhibit weaknesses in certain reasoning and semantic understanding areas.
- Evaluation benchmarks vary in scope and may not fully capture emergent capabilities or safety considerations of LLMs.
- Both automatic and human evaluations are essential, but each has limitations that affect reliability and interpretation.
- There is a need for unified, principled evaluation frameworks that cover general and domain-specific tasks, robustness, and trustworthiness.
- Open-source materials and ongoing benchmark development are crucial for collaborative progress in LLM evaluation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.