Skip to main content
QUICK REVIEW

[論文レビュー] Evaluating Large Language Models on the GMAT: Implications for the Future of Business Education

Vahid Ashrafimoghari, Necdet Gürkan|arXiv (Cornell University)|Jan 2, 2024
Artificial Intelligence in Healthcare and EducationMedicine被引用数 3
ひとこと要約

本研究では、GMATの定量的および言語的セクションにおいて、7つの主要な大規模言語モデル(LLM)を評価し、GPT-4 Turboが人間の受験者および以前のLLMバージョンを上回ることを示した。モデルは強力なゼロショット推論能力を示しており、ビジネス教育におけるAI駆動型指導および評価の変革的潜在能力を示唆している。

ABSTRACT

The rapid evolution of artificial intelligence (AI), especially in the domain of Large Language Models (LLMs) and generative AI, has opened new avenues for application across various fields, yet its role in business education remains underexplored. This study introduces the first benchmark to assess the performance of seven major LLMs, OpenAI's models (GPT-3.5 Turbo, GPT-4, and GPT-4 Turbo), Google's models (PaLM 2, Gemini 1.0 Pro), and Anthropic's models (Claude 2 and Claude 2.1), on the GMAT, which is a key exam in the admission process for graduate business programs. Our analysis shows that most LLMs outperform human candidates, with GPT-4 Turbo not only outperforming the other models but also surpassing the average scores of graduate students at top business schools. Through a case study, this research examines GPT-4 Turbo's ability to explain answers, evaluate responses, identify errors, tailor instructions, and generate alternative scenarios. The latest LLM versions, GPT-4 Turbo, Claude 2.1, and Gemini 1.0 Pro, show marked improvements in reasoning tasks compared to their predecessors, underscoring their potential for complex problem-solving. While AI's promise in education, assessment, and tutoring is clear, challenges remain. Our study not only sheds light on LLMs' academic potential but also emphasizes the need for careful development and application of AI in education. As AI technology advances, it is imperative to establish frameworks and protocols for AI interaction, verify the accuracy of AI-generated content, ensure worldwide access for diverse learners, and create an educational environment where AI supports human expertise. This research sets the stage for further exploration into the responsible use of AI to enrich educational experiences and improve exam preparation and assessment methods.

研究の動機と目的

  • 最先端のLLMがビジネススクール入学のための主要な標準化試験であるGMATにおいていかにパフォーマンスを発揮するかを評価すること。
  • 大学院レベルのビジネス教育に特徴的な推論中心のタスクにおいて、LLMが人間のパフォーマンスに達するか、あるいはそれを上回ることができるかを調査すること。
  • GMAT準備および広範なビジネス教育における、LLMをパーソナライズドチューターとして活用する教育的潜在可能性を探索すること。
  • 学術的評価および学習環境へのLLM統合に伴う倫理的・技術的・教育的課題を特定すること。

提案手法

  • GMAC(大学院経営適性試験委員会)が提供する無料およびプレミアムのGMAT練習試験を用い、データ漏洩や記憶効果の有無を評価した。
  • 微調整やチェーン・オブ・トゥグラウンズ(思考の道筋)プロンプトを一切使用せず、ベースライン推論能力を評価するためのゼロショットプロンプトを適用した。
  • GPT-4 Turboの事例研究を通じて、回答の説明、回答の評価、誤りの特定、代替シナリオの生成能力を分析した。
  • 上位ビジネススクールの平均人間スコアと比較することで、ベンチマークの妥当性を確立した。
  • AIの限界、バイアス、教育的統合リスクを含む定性的分析を通じて、倫理的および実務的影響を評価した。
Figure 1: The template employed for generating prompts for every multiple-choice question. Elements shown in double braces are substituted with question-specific values.
Figure 1: The template employed for generating prompts for every multiple-choice question. Elements shown in double braces are substituted with question-specific values.

実験結果

リサーチクエスチョン

  • RQ1LLMは、GMATの言語的および定量的推論セクションにおいて、人間の受験者と比べてどのようにパフォーマンスを発揮するか?
  • RQ2ビジネス教育における指導、試験準備、評価にLLMを用いる場合の潜在的な利点と欠点は何か?
  • RQ3GPT-4 TurboのようなLLMは、学習を支援する形で、推論を説明し、誤りを検出し、指導を適応的に変更できる程度はどの程度か?

主な発見

  • GPT-4 TurboはGMATで最高の平均スコアを記録し、上位ビジネススクールの平均人間受験者のパフォーマンスを上回った。
  • 評価されたすべてのLLM、特にGPT-4 Turbo、Claude 2.1、Gemini 1.0 Proは、言語的および定量的推論セクションの両方で人間受験者を上回った。
  • 無料版とプレミアム版のGMAT練習試験の間でパフォーマンスに顕著な差はなく、データ漏洩や記憶効果の影響は顕著ではなかった。
  • GPT-4 Turboは、複雑な推論の説明、誤りの特定、逆説的シナリオの生成といった強力な指導能力を示した。
  • GPT-4 Turbo や Claude 2.1 といった最新のLLMバージョンは、それらの前身バージョンと比較して推論能力に顕著な向上を示しており、推論一般化能力の急速な進歩が裏付けられた。
  • 高いパフォーマンスにもかかわらず、LLMは幻覚、事実誤認、曖昧な言語の誤解に起因するリスクを抱えており、教育現場への導入にあたっては懸念が必要である。
Figure 2: An example of implementation of template shown in from Figure 1 .
Figure 2: An example of implementation of template shown in from Figure 1 .

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。