[Paper Review] Automated Feedback in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses
This study evaluates three models—fine-tuned Mistral-based GOAT, SBERT-Canberra, and zero-shot GPT-4—for automated scoring and feedback on open-ended math problems. GOAT outperforms SBERT-Canberra in scoring accuracy, while GPT-4 generates higher-quality, more detailed feedback preferred by human raters, highlighting a trade-off between scoring precision and feedback quality in LLM-based educational systems.
The effectiveness of feedback in enhancing learning outcomes is well documented within Educational Data Mining (EDM). Various prior research has explored methodologies to enhance the effectiveness of feedback. Recent developments in Large Language Models (LLMs) have extended their utility in enhancing automated feedback systems. This study aims to explore the potential of LLMs in facilitating automated feedback in math education. We examine the effectiveness of LLMs in evaluating student responses by comparing 3 different models: Llama, SBERT-Canberra, and GPT4 model. The evaluation requires the model to provide both a quantitative score and qualitative feedback on the student's responses to open-ended math problems. We employ Mistral, a version of Llama catered to math, and fine-tune this model for evaluating student responses by leveraging a dataset of student responses and teacher-written feedback for middle-school math problems. A similar approach was taken for training the SBERT model as well, while the GPT4 model used a zero-shot learning approach. We evaluate the model's performance in scoring accuracy and the quality of feedback by utilizing judgments from 2 teachers. The teachers utilized a shared rubric in assessing the accuracy and relevance of the generated feedback. We conduct both quantitative and qualitative analyses of the model performance. By offering a detailed comparison of these methods, this study aims to further the ongoing development of automated feedback systems and outlines potential future directions for leveraging generative LLMs to create more personalized learning experiences.
Motivation & Objective
- To evaluate the effectiveness of fine-tuned LLMs in generating accurate scores and meaningful feedback for open-ended math problems.
- To compare the performance of a fine-tuned Mistral-based model (GOAT) against SBERT-Canberra and zero-shot GPT-4 in auto-scoring and feedback generation.
- To assess human rater preferences for feedback quality, relevance, and constructiveness across the three models.
- To investigate the impact of training data quality and consistency on model performance, particularly given variability in teacher-graded responses.
- To identify limitations in current automated feedback systems and guide future improvements in personalization and reliability.
Proposed method
- Fine-tuned a Mistral-based LLM (GOAT) on a dataset of student responses paired with teacher-provided scores and feedback for middle-school math problems.
- Trained a SBERT-Canberra model using the same dataset to serve as a baseline for auto-scoring, leveraging semantic similarity for response evaluation.
- Applied a zero-shot prompting strategy with GPT-4, providing detailed rubrics for scoring and feedback generation without fine-tuning.
- Employed a shared rubric used by two experienced teachers to evaluate scoring accuracy and feedback quality in a blinded, comparative assessment.
- Conducted both quantitative analysis of score predictions and qualitative analysis of feedback content using human evaluators.
- Used teacher consistency data from prior studies to contextualize performance variability and inform model training limitations.
Experimental results
Research questions
- RQ1How does the fine-tuned GOAT model compare to SBERT-Canberra in predicting teacher scores for open-ended student responses?
- RQ2How does the zero-shot GPT-4 model compare to the fine-tuned GOAT model in the auto-scoring task for open-ended math problems?
- RQ3Which model—SBERT-Canberra, GOAT, or GPT-4—is preferred by human raters with teaching experience for feedback quality, relevance, and constructiveness?
- RQ4To what extent does the quality of training data, including inconsistent or low-quality teacher feedback, affect the performance of the fine-tuned GOAT model?
- RQ5How do human rater inconsistencies in scoring and feedback influence the evaluation and calibration of automated feedback systems?
Key findings
- GOAT outperformed SBERT-Canberra in scoring accuracy, identifying patterns in teacher scoring that the latter model missed.
- GPT-4 generated feedback that was rated significantly higher in quality, detail, and helpfulness by human evaluators compared to both GOAT and SBERT-Canberra.
- The feedback from GOAT closely mirrored the quality of the training data, including low-quality, vague, or demotivating messages, indicating it learned from flawed teacher examples.
- Human raters showed substantial variability in feedback assessment, with top teachers achieving only 73% consistency above chance, reflecting the subjectivity inherent in human grading.
- The study found that teacher feedback is often not aligned with published rubrics, and many teachers rely on mental models or contextual factors, complicating automated feedback alignment.
- Despite high scoring performance, GOAT’s feedback quality was limited by the quality of the training data, suggesting that model performance is constrained by data quality and diversity.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.