Skip to main content
QUICK REVIEW

[Paper Review] Evaluating Large Language Models on the GMAT: Implications for the Future of Business Education

Vahid Ashrafimoghari, Necdet Gürkan|arXiv (Cornell University)|Jan 2, 2024
Artificial Intelligence in Healthcare and EducationMedicine3 citations
TL;DR

This study evaluates seven leading large language models (LLMs) on the GMAT’s quantitative and verbal sections, demonstrating that GPT-4 Turbo outperforms both human test-takers and earlier LLM versions. The models show strong zero-shot reasoning ability, suggesting transformative potential for AI-driven tutoring and assessment in business education.

ABSTRACT

The rapid evolution of artificial intelligence (AI), especially in the domain of Large Language Models (LLMs) and generative AI, has opened new avenues for application across various fields, yet its role in business education remains underexplored. This study introduces the first benchmark to assess the performance of seven major LLMs, OpenAI's models (GPT-3.5 Turbo, GPT-4, and GPT-4 Turbo), Google's models (PaLM 2, Gemini 1.0 Pro), and Anthropic's models (Claude 2 and Claude 2.1), on the GMAT, which is a key exam in the admission process for graduate business programs. Our analysis shows that most LLMs outperform human candidates, with GPT-4 Turbo not only outperforming the other models but also surpassing the average scores of graduate students at top business schools. Through a case study, this research examines GPT-4 Turbo's ability to explain answers, evaluate responses, identify errors, tailor instructions, and generate alternative scenarios. The latest LLM versions, GPT-4 Turbo, Claude 2.1, and Gemini 1.0 Pro, show marked improvements in reasoning tasks compared to their predecessors, underscoring their potential for complex problem-solving. While AI's promise in education, assessment, and tutoring is clear, challenges remain. Our study not only sheds light on LLMs' academic potential but also emphasizes the need for careful development and application of AI in education. As AI technology advances, it is imperative to establish frameworks and protocols for AI interaction, verify the accuracy of AI-generated content, ensure worldwide access for diverse learners, and create an educational environment where AI supports human expertise. This research sets the stage for further exploration into the responsible use of AI to enrich educational experiences and improve exam preparation and assessment methods.

Motivation & Objective

  • To assess the performance of state-of-the-art LLMs on the GMAT, a key standardized test for business school admissions.
  • To investigate whether LLMs can match or exceed human performance in reasoning-intensive tasks typical of graduate business education.
  • To explore the pedagogical potential of LLMs as personalized tutors in GMAT preparation and broader business education.
  • To identify ethical, technical, and educational challenges in integrating LLMs into academic assessment and learning environments.

Proposed method

  • Benchmarked seven general-purpose LLMs—GPT-3.5 Turbo, GPT-4, GPT-4 Turbo, Claude 2, Claude 2.1, PaLM 2, and Gemini 1.0 Pro—on the GMAT’s quantitative and verbal reasoning sections.
  • Used both free and premium GMAT practice exams from the Graduate Management Admission Council (GMAC) to assess for data leakage or memorization effects.
  • Applied zero-shot prompting without fine-tuning or chain-of-thought prompting to evaluate baseline reasoning capabilities.
  • Conducted a case study on GPT-4 Turbo to analyze its ability to explain answers, evaluate responses, identify errors, and generate alternative scenarios.
  • Compared model performance against average human scores from top business schools to establish benchmark relevance.
  • Evaluated ethical and practical implications through qualitative analysis of AI limitations, bias, and educational integration risks.
Figure 1: The template employed for generating prompts for every multiple-choice question. Elements shown in double braces are substituted with question-specific values.
Figure 1: The template employed for generating prompts for every multiple-choice question. Elements shown in double braces are substituted with question-specific values.

Experimental results

Research questions

  • RQ1How do LLMs compare to human candidates in terms of performance on the GMAT’s verbal and quantitative reasoning sections?
  • RQ2What are the potential benefits and drawbacks of using LLMs for tutoring, exam preparation, and assessment in business education?
  • RQ3To what extent can LLMs like GPT-4 Turbo explain reasoning, detect errors, and adapt instruction in a way that supports learning?

Key findings

  • GPT-4 Turbo achieved the highest average score on the GMAT, surpassing the average performance of human test-takers at top business schools.
  • All evaluated LLMs, especially GPT-4 Turbo, Claude 2.1, and Gemini 1.0 Pro, outperformed human candidates in both verbal and quantitative reasoning sections.
  • The performance difference between free and premium GMAT practice exams was negligible, indicating no significant data leakage or memorization effects.
  • GPT-4 Turbo demonstrated strong tutoring capabilities, including explaining complex reasoning, identifying errors, and generating counterfactual scenarios.
  • Newer LLM versions (e.g., GPT-4 Turbo, Claude 2.1) showed marked improvements in reasoning over their predecessors, indicating rapid progress in reasoning generalization.
  • Despite high performance, LLMs remain prone to hallucination, factual errors, and misinterpretation of nuanced language, highlighting risks in educational deployment.
Figure 2: An example of implementation of template shown in from Figure 1 .
Figure 2: An example of implementation of template shown in from Figure 1 .

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.