Skip to main content
QUICK REVIEW

[논문 리뷰] Evaluating Large Language Models on the GMAT: Implications for the Future of Business Education

Vahid Ashrafimoghari, Necdet Gürkan|arXiv (Cornell University)|2024. 01. 02.
Artificial Intelligence in Healthcare and EducationMedicine인용 수 3
한 줄 요약

이 연구는 경영대학원 입학을 위한 주요 표준화 시험인 GMAT의 수리 및 독해 영역에서 일곱 가지 주요 대규모 언어 모델(Large Language Models, LLMs)의 성능을 평가하며, GPT-4 Turbo가 인간 시험 수험생과 이전 LLM 버전을 모두 능가하는 것으로 나타났다. 모델들은 강력한 제로샷 추론 능력을 보이며, 비즈니스 교육 분야에서 인공지능 기반의 개인 맞춤형 지도 및 평가의 변혁적 잠재력을 시사한다.

ABSTRACT

The rapid evolution of artificial intelligence (AI), especially in the domain of Large Language Models (LLMs) and generative AI, has opened new avenues for application across various fields, yet its role in business education remains underexplored. This study introduces the first benchmark to assess the performance of seven major LLMs, OpenAI's models (GPT-3.5 Turbo, GPT-4, and GPT-4 Turbo), Google's models (PaLM 2, Gemini 1.0 Pro), and Anthropic's models (Claude 2 and Claude 2.1), on the GMAT, which is a key exam in the admission process for graduate business programs. Our analysis shows that most LLMs outperform human candidates, with GPT-4 Turbo not only outperforming the other models but also surpassing the average scores of graduate students at top business schools. Through a case study, this research examines GPT-4 Turbo's ability to explain answers, evaluate responses, identify errors, tailor instructions, and generate alternative scenarios. The latest LLM versions, GPT-4 Turbo, Claude 2.1, and Gemini 1.0 Pro, show marked improvements in reasoning tasks compared to their predecessors, underscoring their potential for complex problem-solving. While AI's promise in education, assessment, and tutoring is clear, challenges remain. Our study not only sheds light on LLMs' academic potential but also emphasizes the need for careful development and application of AI in education. As AI technology advances, it is imperative to establish frameworks and protocols for AI interaction, verify the accuracy of AI-generated content, ensure worldwide access for diverse learners, and create an educational environment where AI supports human expertise. This research sets the stage for further exploration into the responsible use of AI to enrich educational experiences and improve exam preparation and assessment methods.

연구 동기 및 목표

  • 최신 기술의 LLM이 경영대학원 입학을 위한 주요 표준화 시험인 GMAT에서 어떻게 성능을 보이는지 평가하기 위해.
  • 경영대학원 교육에서 일반적인 추론 중심 과제를 수행할 때 LLM이 인간 수준의 성능을 따라하거나 뛰어넘을 수 있는지 조사하기 위해.
  • GMAT 준비 및 더 넓은 비즈니스 교육 분야에서 LLM을 개인 맞춤형 지도자로 활용할 수 있는 교육적 잠재력 탐색하기 위해.
  • LLM을 학술 평가 및 학습 환경에 통합함에 있어 발생할 수 있는 윤리적, 기술적, 교육적 과제 규명하기 위해.

제안 방법

  • 일반 목적의 일곱 가지 LLM(GPT-3.5 Turbo, GPT-4, GPT-4 Turbo, Claude 2, Claude 2.1, PaLM 2, Gemini 1.0 Pro)을 GMAT의 수리 및 독해 추론 영역에서 평가하였다.
  • 데이터 泄露 또는 기억 효과를 검토하기 위해 경영대학원 입학위원회(GMAC)에서 제공하는 무료 및 프리미엄 GMAT 연습 시험을 사용하였다.
  • 세부 조정 또는 사슬 추론 프롬프트 없이 제로샷 프롬프트를 적용하여 기본 추론 능력을 평가하였다.
  • GPT-4 Turbo의 사례 연구를 수행하여 그 모델이 답을 설명하고, 응답을 평가하며, 오류를 식별하고, 대체 시나리오를 생성하는 능력을 분석하였다.
  • 상위 비즈니스 스쿨의 평균 인간 시험 성적과 비교하여 기준 성능의 관련성을 확보하였다.
  • 인공지능의 한계, 편향 및 교육 통합 리스크를 정성적으로 분석하여 윤리적 및 실용적 영향을 평가하였다.
Figure 1: The template employed for generating prompts for every multiple-choice question. Elements shown in double braces are substituted with question-specific values.
Figure 1: The template employed for generating prompts for every multiple-choice question. Elements shown in double braces are substituted with question-specific values.

실험 결과

연구 질문

  • RQ1LLM은 GMAT의 독해 및 수리 추론 영역에서 인간 시험 수험생과 비교해 어떤 성능을 보이는가?
  • RQ2비즈니스 교육 분야에서 지도, 시험 준비 및 평가에 LLM을 사용할 경우 잠재적인 이점과 단점은 무엇인가?
  • RQ3GPT-4 Turbo와 같은 LLM은 학습 지원을 위해 추론을 설명하고 오류를 탐지하며 지도 방식을 어떻게 적응시킬 수 있는가?

주요 결과

  • GPT-4 Turbo는 GMAT에서 가장 높은 평균 점수를 기록하여 상위 비즈니스 스쿨의 평균 인간 시험 수험생 성적을 초월하였다.
  • 평가된 모든 LLM, 특히 GPT-4 Turbo, Claude 2.1, Gemini 1.0 Pro는 독해 및 수리 추론 영역에서 인간 수험생을 모두 뛰어넘는 성과를 보였다.
  • 무료 및 프리미엄 GMAT 연습 시험 간 성과 차이가 미미하여 데이터 泄露 또는 기억 효과가 유의미하게 영향을 주지 않았다.
  • GPT-4 Turbo는 복잡한 추론을 설명하고 오류를 식별하며 반대 상황을 생성하는 강력한 지도 능력을 보였다.
  • 최신 LLM 버전(GPT-4 Turbo, Claude 2.1)은 이전 버전 대비 추론 능력에서 뚜렷한 향상을 보이며 추론 일반화 능력 향상의 급속한 진전을 보였다.
  • 높은 성능에도 불구하고 LLM은 환상, 사실 오류 및 미묘한 언어의 오해를 일으킬 수 있어 교육 현장 적용에 리스크가 있음을 시사한다.
Figure 2: An example of implementation of template shown in from Figure 1 .
Figure 2: An example of implementation of template shown in from Figure 1 .

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.