[论文解读] Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
本研究评估了在德国教育语境下,用于多维作文评分的开源与闭源大型语言模型(LLMs),并将其评分与37名教师在10项标准下的评分进行比较。新型o1模型在评分一致性(ICC = .80)和与人工评分的相关性(Spearman等级相关系数 r = .74)方面表现优异,优于其他模型,尤其在语言相关标准上表现突出,但其整体评分偏高,表明在内容质量评估方面仍需优化。
The manual assessment and grading of student writing is a time-consuming yet critical task for teachers. Recent developments in generative AI, such as large language models, offer potential solutions to facilitate essay-scoring tasks for teachers. In our study, we evaluate the performance and reliability of both open-source and closed-source LLMs in assessing German student essays, comparing their evaluations to those of 37 teachers across 10 pre-defined criteria (i.e., plot logic, expression). A corpus of 20 real-world essays from Year 7 and 8 students was analyzed using five LLMs: GPT-3.5, GPT-4, o1, LLaMA 3-70B, and Mixtral 8x7B, aiming to provide in-depth insights into LLMs' scoring capabilities. Closed-source GPT models outperform open-source models in both internal consistency and alignment with human ratings, particularly excelling in language-related criteria. The novel o1 model outperforms all other LLMs, achieving Spearman's $r = .74$ with human assessments in the overall score, and an internal consistency of $ICC=.80$. These findings indicate that LLM-based assessment can be a useful tool to reduce teacher workload by supporting the evaluation of essays, especially with regard to language-related criteria. However, due to their tendency for higher scores, the models require further refinement to better capture aspects of content quality.
研究动机与目标
- 评估大型语言模型(LLMs)在多维度评分德国学生作文中的可靠性与有效性。
- 比较开源与闭源LLMs在10项预设标准下与人类教师评分的性能表现。
- 评估LLMs在教育作文评分中与人类评估的一致性及其内部一致性。
- 识别LLMs在支持教师评分方面的优势与局限,特别是语言与内容质量评估的差异。
- 为开发能减轻教师工作负担同时保持评估质量的AI辅助作文评分工具提供建议。
提出的方法
- 使用五种LLMs(GPT-3.5、GPT-4、o1、LLaMA 3-70B 和 Mixtral 8x7B)对20篇来自七年级和八年级的真实德语学生作文进行评分。
- 由37名教师依据10项预设标准(包括情节逻辑、表达方式和语言准确性)对作文进行评分。
- LLM评分未经过提示工程处理,以确保一致性并隔离模型行为。
- 采用Spearman等级相关系数(r)和组内相关系数(ICC)等统计指标评估评分一致性与内部一致性。
- 对比分析主要聚焦于闭源模型(如GPT-4、o1)与开源模型(如LLaMA 3、Mixtral)在性能上的差异。
- 通过分析多次LLM运行结果的方差,评估自动化评分的可靠性与稳健性。

实验结果
研究问题
- RQ1在德国作文评分的10项多维标准下,闭源与开源LLMs在与人类教师评分的一致性方面表现如何?
- RQ2LLM生成评分的内部一致性如何?不同模型之间的差异为何?
- RQ3哪一LLM在复制人类教师评分方面表现最佳,特别是在语言相关与内容相关标准方面?
- RQ4LLMs在整体评分上是否表现出对更高分的倾向?这种倾向如何影响其在教育环境中的可靠性?
- RQ5如何优化LLMs以更好地评估内容质量,同时保持与人类评价标准的一致性?
主要发现
- o1模型与人类评分的相关性最高(Spearman等级相关系数 r = .74),内部一致性最强(ICC = .80),优于所有其他模型。
- 闭源模型,尤其是GPT-4和o1,在与人类评分的一致性及内部一致性方面显著优于开源模型。
- 开源模型如LLaMA 3-70B和Mixtral 8x7B表现出较低的方差,但与人类评分的相关性微弱,表明其在多维评估中可靠性较差。
- LLMs在语言相关标准(如表达、语法)上的表现优于内容相关标准(如情节逻辑、论证能力)。
- o1模型的表现表明,新一代经过推理优化的LLMs能够高度模拟人类评分,但其仍倾向于给出更高的整体评分。
- 本研究强调,需进一步优化LLMs以提升内容质量评估能力,并减少评分偏差,以确保与人类标准的一致性。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。