Skip to main content
QUICK REVIEW

[论文解读] ChatGPT & Mechanical Engineering: Examining performance on the FE Mechanical Engineering and Undergraduate Exams

Matthew Frenkel, Hebah Emara|arXiv (Cornell University)|Sep 26, 2023
Artificial Intelligence in Healthcare and EducationMedicine被引用 3
一句话总结

本研究评估了ChatGPT在机械工程考试中的表现,包括FE Mechanical考试以及本科高年级水平的考试。使用GPT-3.5(免费版)和GPT-4(付费版),研究发现GPT-4的准确率为76%,而GPT-3.5仅为51%。然而,由于仅依赖文本输入且对错误答案过于自信,两者的表现均受限,若无专家监督,难以在实际场景中应用。

ABSTRACT

The launch of ChatGPT at the end of 2022 generated large interest into possible applications of artificial intelligence in STEM education and among STEM professions. As a result many questions surrounding the capabilities of generative AI tools inside and outside of the classroom have been raised and are starting to be explored. This study examines the capabilities of ChatGPT within the discipline of mechanical engineering. It aims to examine use cases and pitfalls of such a technology in the classroom and professional settings. ChatGPT was presented with a set of questions from junior and senior level mechanical engineering exams provided at a large private university, as well as a set of practice questions for the Fundamentals of Engineering Exam (FE) in Mechanical Engineering. The responses of two ChatGPT models, one free to use and one paid subscription, were analyzed. The paper found that the subscription model (GPT-4) greatly outperformed the free version (GPT-3.5), achieving 76% correct vs 51% correct, but the limitation of text only input on both models makes neither likely to pass the FE exam. The results confirm findings in the literature with regards to types of errors and pitfalls made by ChatGPT. It was found that due to its inconsistency and a tendency to confidently produce incorrect answers the tool is best suited for users with expert knowledge.

研究动机与目标

  • 评估ChatGPT模型在标准化考试和学术性机械工程考试中的表现。
  • 比较GPT-3.5(免费版)和GPT-4(付费版)在解决机械工程问题方面的能力。
  • 识别生成式AI在STEM应用中输出结果的常见错误模式与可靠性问题。
  • 评估ChatGPT在学术和专业机械工程环境中的适用性。

提出的方法

  • 向一所大型私立大学的本科高年级机械工程课程中的一组题目同时提供给GPT-3.5和GPT-4。
  • 收集两个模型在相同题目集上的回答,包括选择题和问题求解题。
  • 根据标准答案评估回答的准确性和一致性。
  • 分析错误类型,如事实性错误、逻辑缺陷以及对错误答案的过度自信。
  • 采用定性和定量分析方法,比较不同题型和难度级别下模型的表现。
  • 开展对仅限文本输入的局限性及其对FE Mechanical考试准备程度影响的对比分析。

实验结果

研究问题

  • RQ1ChatGPT(GPT-3.5和GPT-4)在解答本科机械工程考试题目时的准确率如何?
  • RQ2ChatGPT在解决机械工程问题时常见的错误类型有哪些?
  • RQ3GPT-4在机械工程考试题目上的表现与GPT-3.5相比如何?
  • RQ4在学术或专业机械工程情境中,在多大程度上可以信任ChatGPT的回答?
  • RQ5仅限文本输入在评估ChatGPT对FE Mechanical等标准化考试准备程度方面存在哪些局限性?

主要发现

  • GPT-4在所测试的机械工程考试题目中达到了76%的准确率,显著优于GPT-3.5。
  • GPT-3.5的准确率为51%,表明其在复杂工程问题求解中的可靠性有限。
  • 两个模型均表现出在多步骤或依赖上下文的问题中,对错误答案过度自信的倾向,尤其在复杂情境下。
  • 缺乏多模态输入(如图表、图像形式的公式)严重限制了模型表现,尤其是在FE Mechanical考试中。
  • 错误模式包括对问题情境的误解、公式应用错误以及推理过程中的逻辑不一致。
  • 尽管回答中表现出高度自信,但由于仅限文本输入的限制以及错误传播问题,两个模型均不太可能通过FE Mechanical考试。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。