[논문 리뷰] A Survey on Evaluation of Large Language Models
이 논문은 무엇을 평가할지, 어디에서 평가할지, 어떻게 평가할지에 걸쳐 대형 언어 모델(LLM)의 평가 방법을 종합적으로 조사하며, 과제, 벤치마크, 도전과제에 주목한다.
Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate, where to evaluate, and how to evaluate. Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, educations, natural and social sciences, agent applications, and other areas. Secondly, we answer the `where' and `how' questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey.
연구 동기 및 목표
- 자연어 처리(NLP), 추론, 윤리, 교육, 과학 및 응용 분야에 걸친 LLM의 기존 평가 과제를 요약한다.
- LLM 성능 평가에 사용되는 평가 데이터셋과 벤치마크를 분석한다.
- 자동 평가와 인간 평가를 포함한 평가 방법론을 논의하고 강점과 한계를 식별한다.
- 원칙에 기반한, 강건하고 포괄적인 LLM 평가를 위한 주요 도전과제와 향후 방향을 강조한다.
제안 방법
- LLM 평가를 무엇을 평가할지, 어디에서 평가할지, 어떻게 평가할지의 세 가지 차원으로 분류한다.
- NLP 과제(NLU, NLG, 추론, 다국어, 사실성) 및 기타 영역(의료, 윤리, 사회과학, 에이전트 응용)을 검토한다.
- LLM 평가에 사용되는 일반 및 구체적 벤치마크와 데이터셋을 목록화한다(예: GLUE, MMLU, BIG-bench 등).
- 자동 평가와 인간 평가 접근법의 차이와 LLM 평가에서의 역할을 검토한다.
- LLM 평가를 위한 주요 도전과제 및 오픈 소스 자원(Open repository)을 논의한다(오픈 리포지토리).

실험 결과
연구 질문
- RQ1LLM을 평가하는 데 사용되는 평가 과제는 무엇이며 그것들이 강점과 약점에 대해 무엇을 드러내는가?
- RQ2어디에서(어떤 데이터셋과 벤치마크에서) LLM이 평가되며, 어떤 벤치마크가 그들의 능력을 잘 포착하는가?
- RQ3LLM은 어떻게 평가되는가(자동 vs. 인간, 프로토콜 설계), 그리고 현재 평가 방법의 한계는 무엇인가?
- RQ4LLM 평가의 주요 도전과제와 향후 방향은 무엇인가?
- RQ5더 강건하고 신뢰할 수 있는 LLM 개발을 가이드하기 위해 어떤 통찰을 얻을 수 있는가?
주요 결과
- LLMs는 많은 NLP 작업에서 강력한 성능을 보이지만 특정 추론 및 의미 이해 영역에서 약점을 보인다.
- 평가 벤치마크는 범위가 다양하며 LLM의 창발적 능력이나 안전성 고려를 완전히 포착하지 못할 수 있다.
- 자동 평가와 인간 평가는 모두 필수적이지만 각각 신뢰성 및 해석에 영향을 주는 한계가 있다.
- 일반 및 도메인 특정 과제를 포괄하고, 강건성 및 신뢰성을 다루는 통합적이고 원칙에 기반한 평가 프레임워크가 필요하다.
- 오픈 소스 자료와 지속적인 벤치마크 개발은 LLM 평가의 공동 발전에 매우 중요하다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.