[Paper Review] ChatGPT & Mechanical Engineering: Examining performance on the FE Mechanical Engineering and Undergraduate Exams
This study evaluates ChatGPT's performance on mechanical engineering exams, including the FE Mechanical and undergraduate junior/senior-level exams. Using GPT-3.5 (free) and GPT-4 (paid), it finds GPT-4 achieved 76% accuracy versus GPT-3.5’s 51%, but both fail due to text-only input and overconfidence in incorrect answers, limiting real-world use without expert oversight.
The launch of ChatGPT at the end of 2022 generated large interest into possible applications of artificial intelligence in STEM education and among STEM professions. As a result many questions surrounding the capabilities of generative AI tools inside and outside of the classroom have been raised and are starting to be explored. This study examines the capabilities of ChatGPT within the discipline of mechanical engineering. It aims to examine use cases and pitfalls of such a technology in the classroom and professional settings. ChatGPT was presented with a set of questions from junior and senior level mechanical engineering exams provided at a large private university, as well as a set of practice questions for the Fundamentals of Engineering Exam (FE) in Mechanical Engineering. The responses of two ChatGPT models, one free to use and one paid subscription, were analyzed. The paper found that the subscription model (GPT-4) greatly outperformed the free version (GPT-3.5), achieving 76% correct vs 51% correct, but the limitation of text only input on both models makes neither likely to pass the FE exam. The results confirm findings in the literature with regards to types of errors and pitfalls made by ChatGPT. It was found that due to its inconsistency and a tendency to confidently produce incorrect answers the tool is best suited for users with expert knowledge.
Motivation & Objective
- To assess the performance of ChatGPT models on standardized and academic mechanical engineering exams.
- To compare the capabilities of GPT-3.5 (free) and GPT-4 (paid) in solving mechanical engineering problems.
- To identify common error patterns and reliability issues in generative AI outputs for STEM applications.
- To evaluate the suitability of ChatGPT for use in academic and professional mechanical engineering settings.
Proposed method
- Administered a set of questions from junior and senior-level mechanical engineering courses at a large private university to both GPT-3.5 and GPT-4.
- Collected responses from both models on the same set of questions, including multiple-choice and problem-solving items.
- Evaluated responses against answer keys to determine accuracy and consistency.
- Analyzed error types, such as factual inaccuracies, logical flaws, and overconfidence in incorrect answers.
- Used qualitative and quantitative analysis to compare model performance across question types and difficulty levels.
- Conducted a comparative analysis of text-only input limitations and their impact on exam readiness.
Experimental results
Research questions
- RQ1How accurately does ChatGPT (GPT-3.5 and GPT-4) solve problems from undergraduate mechanical engineering exams?
- RQ2What types of errors does ChatGPT commonly make when solving mechanical engineering problems?
- RQ3How does the performance of GPT-4 compare to GPT-3.5 on mechanical engineering exam questions?
- RQ4To what extent can ChatGPT’s responses be trusted in academic or professional mechanical engineering contexts?
- RQ5What are the limitations of text-only input in assessing ChatGPT’s readiness for standardized exams like the FE Mechanical?
Key findings
- GPT-4 achieved 76% accuracy on the tested mechanical engineering exam questions, significantly outperforming GPT-3.5.
- GPT-3.5 achieved 51% accuracy, indicating limited reliability for complex engineering problem-solving.
- Both models exhibited a tendency to confidently produce incorrect answers, especially in multi-step or context-sensitive problems.
- The lack of multimodal input (e.g., diagrams, equations in image form) severely limited performance, particularly on the FE Mechanical exam.
- Error patterns included misinterpretation of problem context, incorrect application of formulas, and logical inconsistencies in reasoning.
- Despite high confidence in responses, neither model is likely to pass the FE Mechanical exam due to text-only input constraints and error propagation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.