Skip to main content
QUICK REVIEW

[論文レビュー] ChatGPT & Mechanical Engineering: Examining performance on the FE Mechanical Engineering and Undergraduate Exams

Matthew Frenkel, Hebah Emara|arXiv (Cornell University)|Sep 26, 2023
Artificial Intelligence in Healthcare and EducationMedicine被引用数 3
ひとこと要約

本研究では、機械工学の試験、特にFE Mechanicalおよび大学3年次・4年次レベルの試験において、ChatGPTのパフォーマンスを評価している。GPT-3.5(無料)とGPT-4(有料)を用いて分析した結果、GPT-4は76%の正答率を示した一方、GPT-3.5は51%にとどまり、両者ともテキストのみの入力と誤った回答に対する過信のため、専門的・実務的利用には熟練した監視が不可欠であることが判明した。

ABSTRACT

The launch of ChatGPT at the end of 2022 generated large interest into possible applications of artificial intelligence in STEM education and among STEM professions. As a result many questions surrounding the capabilities of generative AI tools inside and outside of the classroom have been raised and are starting to be explored. This study examines the capabilities of ChatGPT within the discipline of mechanical engineering. It aims to examine use cases and pitfalls of such a technology in the classroom and professional settings. ChatGPT was presented with a set of questions from junior and senior level mechanical engineering exams provided at a large private university, as well as a set of practice questions for the Fundamentals of Engineering Exam (FE) in Mechanical Engineering. The responses of two ChatGPT models, one free to use and one paid subscription, were analyzed. The paper found that the subscription model (GPT-4) greatly outperformed the free version (GPT-3.5), achieving 76% correct vs 51% correct, but the limitation of text only input on both models makes neither likely to pass the FE exam. The results confirm findings in the literature with regards to types of errors and pitfalls made by ChatGPT. It was found that due to its inconsistency and a tendency to confidently produce incorrect answers the tool is best suited for users with expert knowledge.

研究の動機と目的

  • ChatGPTモデルが標準化されたおよび学術的な機械工学試験でどのように機能するかを評価すること。
  • GPT-3.5(無料)とGPT-4(有料)が機械工学的問題を解く能力を比較すること。
  • STEM応用における生成型AI出力の一般的な誤りパターンと信頼性の問題を特定すること。
  • ChatGPTが学術的および職業的機械工学の文脈で使用可能かどうかを評価すること。

提案手法

  • 大規模な私立大学の機械工学専攻の3年次・4年次レベルの試験問題を、GPT-3.5およびGPT-4に提示した。
  • 複数選択および問題解決形式の両方の問題に対して、両モデルの回答を同一のセットで収集した。
  • 正答キーと照合して、正答率と一貫性を評価した。
  • 事実誤認、論理的欠陥、誤った回答に対する過信といった誤りタイプを分析した。
  • 質的および定量的分析を用いて、問題の種別および難易度ごとのモデルパフォーマンスを比較した。
  • テキストのみの入力の制限とそのFE Mechanical試験への準備状況への影響を比較分析した。

実験結果

リサーチクエスチョン

  • RQ1ChatGPT(GPT-3.5およびGPT-4)は、大学レベルの機械工学試験問題をどの程度正確に解けるか?
  • RQ2ChatGPTは機械工学的問題を解く際に一般的にどのような誤りを犯すか?
  • RQ3GPT-4のパフォーマンスはGPT-3.5と比べてどの程度優れているか?
  • RQ4学術的または職業的機械工学の文脈において、ChatGPTの回答をどの程度信頼できるか?
  • RQ5テキストのみの入力は、FE Mechanicalのような標準化試験へのChatGPTの準備状況にどの程度制限をもたらすか?

主な発見

  • GPT-4は、テストされた機械工学試験問題に対して76%の正答率を示し、GPT-3.5を顕著に上回った。
  • GPT-3.5は51%の正答率を示し、複雑な工学的問題解決において限られた信頼性しか示さなかった。
  • 両モデルとも、特に複数ステップまたは文脈依存の問題では、誤った回答を自信を持って提示する傾向を示した。
  • 図や数式を画像形式で提供するマルチモーダル入力がないことが、特にFE Mechanical試験においてパフォーマンスを著しく制限した。
  • 誤りのパターンには、問題文の文脈の誤解、公式の誤った適用、推論における論理的不整合が含まれた。
  • 高い自信を示したにもかかわらず、テキストのみの入力制限と誤りの累積的拡大のため、両モデルともFE Mechanical試験に合格する可能性は極めて低い。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。