Skip to main content
QUICK REVIEW

[论文解读] Evaluating Large Language Models on the GMAT: Implications for the Future of Business Education

Vahid Ashrafimoghari, Necdet Gürkan|arXiv (Cornell University)|Jan 2, 2024
Artificial Intelligence in Healthcare and EducationMedicine被引用 3
一句话总结

本研究评估了七种领先的大型语言模型(LLMs)在GMAT定量和语文部分的表现,表明GPT-4 Turbo在表现上超越了人类考生及早期的LLM版本。这些模型展现出强大的零样本推理能力,表明其在商业教育中驱动个性化辅导与评估方面具有变革性潜力。

ABSTRACT

The rapid evolution of artificial intelligence (AI), especially in the domain of Large Language Models (LLMs) and generative AI, has opened new avenues for application across various fields, yet its role in business education remains underexplored. This study introduces the first benchmark to assess the performance of seven major LLMs, OpenAI's models (GPT-3.5 Turbo, GPT-4, and GPT-4 Turbo), Google's models (PaLM 2, Gemini 1.0 Pro), and Anthropic's models (Claude 2 and Claude 2.1), on the GMAT, which is a key exam in the admission process for graduate business programs. Our analysis shows that most LLMs outperform human candidates, with GPT-4 Turbo not only outperforming the other models but also surpassing the average scores of graduate students at top business schools. Through a case study, this research examines GPT-4 Turbo's ability to explain answers, evaluate responses, identify errors, tailor instructions, and generate alternative scenarios. The latest LLM versions, GPT-4 Turbo, Claude 2.1, and Gemini 1.0 Pro, show marked improvements in reasoning tasks compared to their predecessors, underscoring their potential for complex problem-solving. While AI's promise in education, assessment, and tutoring is clear, challenges remain. Our study not only sheds light on LLMs' academic potential but also emphasizes the need for careful development and application of AI in education. As AI technology advances, it is imperative to establish frameworks and protocols for AI interaction, verify the accuracy of AI-generated content, ensure worldwide access for diverse learners, and create an educational environment where AI supports human expertise. This research sets the stage for further exploration into the responsible use of AI to enrich educational experiences and improve exam preparation and assessment methods.

研究动机与目标

  • 评估最先进的LLMs在GMAT上的表现,该考试是商学院入学的关键标准化测试。
  • 探究LLMs是否能在典型研究生商业教育中所需的推理密集型任务中达到或超越人类表现。
  • 探索LLMs作为GMAT备考及更广泛商业教育中个性化导师的教育潜力。
  • 识别将LLMs整合到学术评估与学习环境时面临的技术、伦理与教育挑战。

提出的方法

  • 在GMAT的定量与语文推理部分,对七种通用型LLMs——GPT-3.5 Turbo、GPT-4、GPT-4 Turbo、Claude 2、Claude 2.1、PaLM 2和Gemini 1.0 Pro——进行基准测试。
  • 使用由管理研究生入学考试委员会(GMAC)提供的免费与付费GMAT练习考试,以评估是否存在数据泄露或记忆效应。
  • 采用零样本提示方法,未进行微调或思维链提示,以评估其基础推理能力。
  • 对GPT-4 Turbo开展案例研究,分析其在解释答案、评估回答、识别错误以及生成替代情景方面的能力。
  • 将模型表现与顶尖商学院人类考生的平均分数进行对比,以确立基准的相关性。
  • 通过定性分析AI的局限性、偏见及教育整合风险,评估伦理与实际影响。
Figure 1: The template employed for generating prompts for every multiple-choice question. Elements shown in double braces are substituted with question-specific values.
Figure 1: The template employed for generating prompts for every multiple-choice question. Elements shown in double braces are substituted with question-specific values.

实验结果

研究问题

  • RQ1LLMs在GMAT语文与定量推理部分的表现与人类考生相比如何?
  • RQ2在商业教育中用于辅导、考试准备与评估时,LLMs的潜在优势与劣势是什么?
  • RQ3LLMs如GPT-4 Turbo在解释推理、检测错误以及以支持学习的方式调整教学方面,其能力达到何种程度?

主要发现

  • GPT-4 Turbo在GMAT上取得了最高平均分,超越了顶尖商学院人类考生的平均水平。
  • 所有评估的LLMs,尤其是GPT-4 Turbo、Claude 2.1和Gemini 1.0 Pro,在语文与定量推理部分的表现均优于人类考生。
  • 免费与付费GMAT练习考试之间的表现差异可忽略不计,表明不存在显著的数据泄露或记忆效应。
  • GPT-4 Turbo展现出强大的辅导能力,包括解释复杂推理、识别错误以及生成反事实情景。
  • 较新版本的LLMs(如GPT-4 Turbo、Claude 2.1)相较于其前代版本在推理能力上表现出显著提升,表明推理泛化能力正在快速进步。
  • 尽管表现优异,LLMs仍易出现幻觉、事实性错误以及对复杂语言细微差别的误解,凸显其在教育环境中部署的风险。
Figure 2: An example of implementation of template shown in from Figure 1 .
Figure 2: An example of implementation of template shown in from Figure 1 .

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。