[论文解读] Automated Feedback in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses
本研究评估了三种模型——微调的基于Mistral的GOAT、SBERT-Canberra和零样本GPT-4——在开放性数学问题自动评分与反馈生成中的表现。GOAT在评分准确性方面优于SBERT-Canberra,而GPT-4生成的反馈在质量上更优、更详细,更受人类评分者青睐,凸显了大语言模型教育系统中评分精度与反馈质量之间的权衡。
The effectiveness of feedback in enhancing learning outcomes is well documented within Educational Data Mining (EDM). Various prior research has explored methodologies to enhance the effectiveness of feedback. Recent developments in Large Language Models (LLMs) have extended their utility in enhancing automated feedback systems. This study aims to explore the potential of LLMs in facilitating automated feedback in math education. We examine the effectiveness of LLMs in evaluating student responses by comparing 3 different models: Llama, SBERT-Canberra, and GPT4 model. The evaluation requires the model to provide both a quantitative score and qualitative feedback on the student's responses to open-ended math problems. We employ Mistral, a version of Llama catered to math, and fine-tune this model for evaluating student responses by leveraging a dataset of student responses and teacher-written feedback for middle-school math problems. A similar approach was taken for training the SBERT model as well, while the GPT4 model used a zero-shot learning approach. We evaluate the model's performance in scoring accuracy and the quality of feedback by utilizing judgments from 2 teachers. The teachers utilized a shared rubric in assessing the accuracy and relevance of the generated feedback. We conduct both quantitative and qualitative analyses of the model performance. By offering a detailed comparison of these methods, this study aims to further the ongoing development of automated feedback systems and outlines potential future directions for leveraging generative LLMs to create more personalized learning experiences.
研究动机与目标
- 评估微调的大语言模型在生成开放性数学问题准确评分与有意义反馈方面的有效性。
- 比较基于Mistral的微调模型(GOAT)与SBERT-Canberra及零样本GPT-4在自动评分与反馈生成中的表现。
- 评估具有教学经验的人类评分者对三种模型在反馈质量、相关性与建设性方面的偏好。
- 探究训练数据质量与一致性对模型性能的影响,特别是考虑到教师评分结果的变异性。
- 识别当前自动反馈系统中的局限性,并为个性化与可靠性方面的未来改进提供指导。
提出的方法
- 在包含学生作答与教师提供评分及反馈的中学数学问题数据集上,对基于Mistral的大语言模型(GOAT)进行微调。
- 使用相同数据集训练SBERT-Canberra模型,作为自动评分的基线,利用语义相似度评估学生作答。
- 对GPT-4采用零样本提示策略,提供详细的评分与反馈生成评分标准,无需微调。
- 采用两名经验丰富的教师共同使用的统一评分标准,在盲评对比评估中衡量评分准确度与反馈质量。
- 对评分预测结果进行定量分析,并通过人类评估者对反馈内容进行定性分析。
- 利用先前研究中教师评分一致性数据,为模型性能波动提供背景信息,并指导模型训练的局限性。
实验结果
研究问题
- RQ1微调的GOAT模型在预测教师对开放性学生作答评分方面,与SBERT-Canberra相比表现如何?
- RQ2零样本GPT-4模型在开放性数学问题自动评分任务中,与微调的GOAT模型相比表现如何?
- RQ3在反馈质量、相关性与建设性方面,人类评分者更偏好哪一模型——SBERT-Canberra、GOAT还是GPT-4?
- RQ4训练数据质量(包括不一致或低质量的教师反馈)在多大程度上影响微调GOAT模型的性能?
- RQ5人类评分者在评分与反馈方面的不一致性,如何影响自动反馈系统的评估与校准?
主要发现
- GOAT在评分准确性方面优于SBERT-Canberra,能够识别出后者所遗漏的教师评分模式。
- 人类评估者对GPT-4生成的反馈在质量、细节与实用性方面评分显著更高,优于GOAT与SBERT-Canberra。
- GOAT生成的反馈质量与训练数据高度一致,包括低质量、模糊或缺乏激励性的反馈,表明其学习了教师提供的有缺陷示例。
- 人类评分者在反馈评估方面表现出显著的不一致性,即使顶尖教师的一致性也仅达到73%(高于随机水平),反映出人类评分固有的主观性。
- 研究发现,教师反馈往往与公开发布的评分标准不一致,许多教师依赖心理模型或情境因素,这使得自动反馈对齐更加复杂。
- 尽管评分表现优异,GOAT的反馈质量仍受限于训练数据的质量,表明模型性能受数据质量与多样性制约。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。