Skip to main content
QUICK REVIEW

[논문 리뷰] A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Hugh Zhang, Jeff Da|arXiv (Cornell University)|2024. 05. 01.
Educational Assessment and Pedagogy인용 수 10
한 줄 요약

이 논문은 GSM1k를 도입하여 GSM8k를 모방한 1250개 문제 벤치마크를 통해 LLM의 학년 수준 산술 수행이 진짜 추론인지 데이터 오염에 의한 것인지 평가하고, 여러 모델 계열에서 상당한 과적합을 발견했지만 프런티어(frontier) 모델에서 강한 일반화를 보였다.

ABSTRACT

Large language models (LLMs) have achieved impressive success on many benchmarks for mathematical reasoning. However, there is growing concern that some of this performance actually reflects dataset contamination, where data closely resembling benchmark questions leaks into the training data, instead of true reasoning ability. To investigate this claim rigorously, we commission Grade School Math 1000 (GSM1k). GSM1k is designed to mirror the style and complexity of the established GSM8k benchmark, the gold standard for measuring elementary mathematical reasoning. We ensure that the two benchmarks are comparable across important metrics such as human solve rates, number of steps in solution, answer magnitude, and more. When evaluating leading open- and closed-source LLMs on GSM1k, we observe accuracy drops of up to 8%, with several families of models showing evidence of systematic overfitting across almost all model sizes. Further analysis suggests a positive relationship (Spearman's r^2 = 0.36) between a model's probability of generating an example from GSM8k and its performance gap between GSM8k and GSM1k, suggesting that some models may have partially memorized GSM8k. Nevertheless, many models, especially those on the frontier, show minimal signs of overfitting, and all models broadly demonstrate generalization to novel math problems guaranteed to not be in their training data.

연구 동기 및 목표

  • GSM8k 스타일 벤치마크가 데이터 오염으로 고생하는지 여부를 인간이 만든 독립적 GSM1k 데이터셋으로 구성해 평가한다.
  • 다양한 모델 패밀리와 크기에 걸쳐 GSM1k와 GSM8k를 비교해 overfitting과 일반화를 정량화한다.

제안 방법

  • GSM1k를 GSM8k의 난이도 분포에 맞춘 1250개의 인간 주석 문제로 구성한다.
  • EleutherAI LM Evaluation Harness의 포크를 사용해 GSM1k에 대해 오픈 소스 및 폐쇄 소스 LLM을 평가한다.
  • 모델 군 간의 GSM8k vs GSM1k 수행 차이를 비교해 과적합을 분석한다.
  • 데이터 오염을 평가하기 위해 GSM8k 사례를 생성할 가능성을 측정한다.
  • 정성적 교훈을 제공하고 오염 외의 과적합 원인에 대해 토론한다.
Figure 1 : Notable models arranged by their drop in performance between GSM8k and GSM1k (lower is worse). We notice that Mistral and Phi top the list of overfit models, with almost 10% drops on GSM1k compared to GSM8k, while models such as Gemini, GPT, and Claude show little to no signs of overfitti
Figure 1 : Notable models arranged by their drop in performance between GSM8k and GSM1k (lower is worse). We notice that Mistral and Phi top the list of overfit models, with almost 10% drops on GSM1k compared to GSM8k, while models such as Gemini, GPT, and Claude show little to no signs of overfitti

실험 결과

연구 질문

  • RQ1GSM1k가 GSM8k에서 감추고 있을 수 있는 과적합을 드러내는가?
  • RQ2크기와 릴리스에 걸쳐 체계적 과적합을 보이는 모델 패밀리는 어떤 것인가?
  • RQ3프런티어 모델은 과적합이 덜하고 새로운 문제에 대한 일반화가 더 나은가?
  • RQ4모델의 GSM8k 데이터를 생성할 가능성과 GSM8k–GSM1k 성능 격차 사이의 관계는 무엇인가?
  • RQ5데이터 오염이 관측된 과적합을 설명하는 데 얼마나 큰 기여를 하는가, 다른 요인은 무엇이 있는가?

주요 결과

  • GSM1k는 여러 모델 패밀리에서 GSM1k 대비 GSM8k에서의 정확도가 최대 13%P 감소하는 현상을 보인다.
  • Mistral 및 Phi 패밀리는 모델 크기 전반에 걸쳐 체계적 과적합을 보이는 반면 프런티어 모델은 과적합이 거의 나타나지 않는다.
  • 프런티어 모델(예: Gemini, GPT, Claude)은 GSM8k와 GSM1k에서 비슷한 성능을 보이며, 일반화가 더 강하거나 오염 방지력이 높은 것으로 해석된다.
  • 모델의 GSM8k 데이터 생성 가능성과 GSM8k–GSM1k 성능 격차 사이에는 Spearman r^2 = 0.32의 양의 관계가 있어 GSM8k 테스트 데이터의 부분적 기억화를 시사한다.
  • 과적합된 모델도 효과적으로 추론하고 새로운 GSM1k 문제를 해결할 수 있어 과적합이 곧 추론 능력을 완전히 제거하지 않는다는 점을 도전한다.
  • 데이터 오염은 과적합의 유일한 설명이 아닐 가능성이 있으며 벤치마크 기반 데이터 수집과 같은 다른 요인도 기여할 수 있다.
Figure 3 : Approximate difficulty distribution of GSM8k train and test sets, measured by number of required steps to solve the problem. GSM1k annotators were instructed to create problems matching the overall distribution of the combined train and test difficulty distribution. The process of estimat
Figure 3 : Approximate difficulty distribution of GSM8k train and test sets, measured by number of required steps to solve the problem. GSM1k annotators were instructed to create problems matching the overall distribution of the combined train and test difficulty distribution. The process of estimat

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.