Skip to main content
QUICK REVIEW

[Paper Review] AI and the FCI: Can ChatGPT Project an Understanding of Introductory Physics?

Colin G. West|ArXiv.org|Mar 2, 2023
Explainable Artificial Intelligence (XAI)37 citations
TL;DR

The paper evaluates two ChatGPT versions (3.5 and 4) using a modified Force Concept Inventory to assess conceptual understanding in introductory physics, finding 3.5 around a typical first-semester student and 4 approaching expert-level performance on mechanics questions.

ABSTRACT

ChatGPT is a groundbreaking ``chatbot"--an AI interface built on a large language model that was trained on an enormous corpus of human text to emulate human conversation. Beyond its ability to converse in a plausible way, it has attracted attention for its ability to competently answer questions from the bar exam and from MBA coursework, and to provide useful assistance in writing computer code. These apparent abilities have prompted discussion of ChatGPT as both a threat to the integrity of higher education and conversely as a powerful teaching tool. In this work we present a preliminary analysis of how two versions of ChatGPT (ChatGPT3.5 and ChatGPT4) fare in the field of first-semester university physics, using a modified version of the Force Concept Inventory (FCI) to assess whether it can give correct responses to conceptual physics questions about kinematics and Newtonian dynamics. We demonstrate that, by some measures, ChatGPT3.5 can match or exceed the median performance of a university student who has completed one semester of college physics, though its performance is notably uneven and the results are nuanced. By these same measures, we find that ChatGPT4's performance is approaching the point of being indistinguishable from that of an expert physicist when it comes to introductory mechanics topics. After the completion of our work we became aware of Ref [1], which preceded us to publication and which completes an extensive analysis of the abilities of ChatGPT3.5 in a physics class, including a different modified version of the FCI. We view this work as confirming that portion of their results, and extending the analysis to ChatGPT4, which shows rapid and notable improvement in most, but not all respects.

Motivation & Objective

  • Assess whether ChatGPT can exhibit conceptual understanding in introductory physics through the FCI.
  • Compare ChatGPT3.5 and ChatGPT4 performance to human students and experts.
  • Explore how prompt engineering and question modification affect model responses.

Proposed method

  • Use a modified, text-only version of the 30-item Force Concept Inventory (FCI) to test ChatGPT.
  • Convert figure-reliant items into text-described prompts to enable processing by ChatGPT3.5 and 4.
  • Administer questions in BASIC and NOVICE prompting styles to evaluate reasoning and stability of responses.
  • Analyze multiple-choice accuracy and qualitative explanations to gauge apparent understanding vs. correct answers.
  • Compare model results to historical student post-test distributions from a large introductory physics course.

Experimental results

Research questions

  • RQ1Can ChatGPT produce correct responses to conceptual kinematics and Newtonian dynamics questions as measured by the FCI?
  • RQ2How do ChatGPT3.5 and ChatGPT4 compare in accuracy and depth of reasoning on introductory physics concepts?
  • RQ3To what extent does prompt framing (BASIC vs NOVICE) and question modification (textual descriptions of figures) affect performance?

Key findings

  • ChatGPT3.5 answered 15 of 23 usable FCI items correctly with BASIC prompting.
  • ChatGPT4 answered 22 of 23 usable FCI items correctly with BASIC prompting, missing item 26 under certain assumptions (air resistance neglected).
  • ChatGPT4’s performance is near the level of an expert physicist for introductory mechanics topics under BASIC prompting.
  • Free-response explanations from ChatGPT3.5 were entirely correct in 10 of 23 cases and broadly correct but with errors in others.
  • ChatGPT3.5 showed substantial weaknesses on spatial-reasoning items involving figures, whereas ChatGPT4 eliminated most of these issues.
  • Results align with prior work showing ChatGPT can display appearance of understanding, and show rapid improvement from 3.5 to 4.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.