Skip to main content
QUICK REVIEW

[论文解读] Performance of Large Language Models in a Computer Science Degree Program

Timothy Krüger, Michael Gref|arXiv (Cornell University)|Jul 24, 2023
Artificial Intelligence in Healthcare and EducationMedicine被引用 3
一句话总结

本研究评估了大型语言模型(LLMs)如 GPT-4.0、ChatGPT-3.5、BingAI 和 LLaMA 变体在计算机科学本科课程中的表现。通过使用课程内容、练习题和历年考试题对模型进行提示,覆盖10个模块,研究发现 GPT-4.0 平均得分达到 80.2%,但仍因数学推理能力薄弱而未能通过学位考核,凸显了尽管在文本类任务中表现优异,但其仍存在关键局限性。

ABSTRACT

Large language models such as ChatGPT-3.5 and GPT-4.0 are ubiquitous and dominate the current discourse. Their transformative capabilities have led to a paradigm shift in how we interact with and utilize (text-based) information. Each day, new possibilities to leverage the capabilities of these models emerge. This paper presents findings on the performance of different large language models in a university of applied sciences' undergraduate computer science degree program. Our primary objective is to assess the effectiveness of these models within the curriculum by employing them as educational aids. By prompting the models with lecture material, exercise tasks, and past exams, we aim to evaluate their proficiency across different computer science domains. We showcase the strong performance of current large language models while highlighting limitations and constraints within the context of such a degree program. We found that ChatGPT-3.5 averaged 79.9% of the total score in 10 tested modules, BingAI achieved 68.4%, and LLaMa, in the 65 billion parameter variant, 20%. Despite these convincing results, even GPT-4.0 would not pass the degree program - due to limitations in mathematical calculations.

研究动机与目标

  • 评估大型语言模型(LLMs)作为计算机科学本科课程教育辅助工具的有效性。
  • 评估 LLM 在涵盖理论、编程和数学成分的多样化计算机科学模块中的表现。
  • 识别主流 LLM(尤其是 GPT-4.0、ChatGPT-3.5、BingAI 和 LLaMA 变体)在学术任务中的优势与局限性。
  • 考察当前 LLM 是否可作为学生学习的有效工具,或是否对学术诚信构成风险。
  • 基于 LLM 的能力与不足,为课程设计与评估策略提供建议。

提出的方法

  • 通过在 10 个计算机科学模块的真实课程材料(包括讲义、练习题和历年考试)上测试 LLM,共收集 40 个数据点。
  • 使用基于实际课程内容的标准化提示,以确保一致性和现实相关性。
  • 在文本推理和数学问题求解任务上评估模型,得分已标准化为总分的百分比。
  • 选取涵盖不同架构和参数规模的模型:GPT-4.0、ChatGPT-3.5、BingAI、LLaMA-65B-Q、LLaMA-7B-Q 和 StableLM-7B。
  • 采用对比评估框架,计算各模型在测试模块中的平均得分。
  • 评估模型在复杂或计算密集型任务中的上下文保持能力、事实准确性与连贯性。

实验结果

研究问题

  • RQ1主流 LLM 在真实本科计算机科学课程材料(包括考试和练习题)上的表现如何?
  • RQ2GPT-4.0、ChatGPT-3.5、BingAI 和较小模型如 LLaMA-65B-Q 在学术计算机科学任务中的关键优势与劣势是什么?
  • RQ3LLM 在多个模块上的表现能否使其通过计算机科学学位课程?
  • RQ4数学推理与计算准确性在技术性计算机科学领域如何影响 LLM 的表现?
  • RQ5LLM 能否作为高等教育中的有效教育辅助工具,还是其局限性削弱了其实用性?

主要发现

  • GPT-4.0 在六个测试模块中平均得分为 80.2%,在文本类和概念性计算机科学任务中表现强劲。
  • ChatGPT-3.5 在十个模块中得分为 79.9%,通过九个模块,表明在非数学领域具有高度可靠性。
  • BingAI 在十个模块中得分为 68.4%,失败两个模块,尤其在无法搜索或语境复杂的题目上表现不佳。
  • LLaMA-65B-Q 在两个模块中平均得分为 20.0%,表现出有限能力,且未能正确解决任何任务。
  • LLaMA-7B-Q 在六个模块中仅得 12.3%,所有任务均失败,常出现上下文丢失或生成无关内容。
  • StableLM-7B 得分最低,仅为 10.8%,未能正确回答任何问题,且频繁在商业语境中误解问题。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。