[Paper Review] Can an AI-tool grade assignments in an introductory physics course?
This study explores an AI-assisted workflow for grading handwritten physics problem solutions using GPT-4, simulating OCR conversion to LaTeX and AI-based feedback. While GPT-4 provides plausible, formative feedback, it lacks reliability for high-stakes summative assessment, making it best suited for initial screening and flagging complex solution paths for human review.
Problem solving is an integral part of any physics curriculum, and most physics instructors would likely agree that the associated learner competencies are best assessed by considering the solution path: not only the final solution matters, but also how the learner arrived there. Unfortunately, providing meaningful feedback on written derivations is much more labor and resource intensive than only grading the outcome: currently, the latter can be done by computer, while the former involves handwritten solutions that need to be graded by humans. This exploratory study proposes an AI-assisted workflow for grading written physics-problem solutions, and it evaluates the viability of the actual grading step using GPT-4. It is found that the AI-tool is capable of providing feedback that can be helpful in formative assessment scenarios, but that for summative scenarios, particularly those that are high-stakes, it should only be used for an initial round of grading that sorts and flags solution approaches.
Motivation & Objective
- To investigate whether GPT-4 can provide meaningful feedback on handwritten physics problem solutions.
- To assess the feasibility of an AI-assisted grading workflow integrating OCR and LLMs for introductory physics courses.
- To determine the suitability of AI grading for formative versus summative assessment contexts.
- To identify limitations of current LLMs in evaluating solution paths and derivations with high reliability.
Proposed method
- The study uses GPT-4 to grade 25 unique handwritten-style physics solutions generated via prompt repetition.
- Solutions were simulated as LaTeX documents after OCR processing, though actual OCR was not tested due to API access restrictions.
- Multiple independent grading rounds were conducted to assess consistency and detect divergent feedback.
- Feedback was analyzed for plausibility, completeness, and alignment with established problem-solving rubrics.
- The workflow integrates scanning, OCR-to-LaTeX conversion, and AI grading, with potential for human arbitration on flagged cases.
- The approach is designed for offline, paper-based assessments to maintain exam integrity.
Experimental results
Research questions
- RQ1Can GPT-4 provide feedback on handwritten physics solutions that is both plausible and useful for formative assessment?
- RQ2How consistent is GPT-4’s grading across multiple independent rounds for the same solution?
- RQ3To what extent does GPT-4’s feedback align with established problem-solving rubrics in physics education?
- RQ4What are the limitations of using GPT-4 for high-stakes summative assessment in introductory physics courses?
- RQ5Can AI grading effectively flag non-standard or divergent solution paths for subsequent human review?
Key findings
- GPT-4 generated plausible narrative feedback on physics solutions, demonstrating potential for use in formative assessment.
- The AI’s feedback was inconsistent across multiple grading rounds, indicating variability in its evaluation of the same solution.
- GPT-4 frequently failed to detect or correctly evaluate non-standard or flawed solution paths, particularly when they deviated from conventional approaches.
- The system’s feedback, while often reasonable, lacked the reliability required for high-stakes summative assessment.
- GPT-4 is best suited for an initial screening phase that flags complex or divergent solutions for human review, rather than full summative grading.
- The study confirms that current LLMs like GPT-4 are not yet reliable for replacing human grading in high-stakes contexts, especially when solution paths are critical to assessment.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.