[Paper Review] Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
This study evaluates open-source and closed-source large language models (LLMs) for multidimensional essay scoring in German education, comparing their ratings to those of 37 teachers across 10 criteria. The novel o1 model achieves high reliability (ICC = .80) and strong correlation (Spearman’s r = .74) with human ratings, outperforming other models, especially in language-related criteria, though it tends to assign higher overall scores, indicating a need for refinement in content quality assessment.
The manual assessment and grading of student writing is a time-consuming yet critical task for teachers. Recent developments in generative AI, such as large language models, offer potential solutions to facilitate essay-scoring tasks for teachers. In our study, we evaluate the performance and reliability of both open-source and closed-source LLMs in assessing German student essays, comparing their evaluations to those of 37 teachers across 10 pre-defined criteria (i.e., plot logic, expression). A corpus of 20 real-world essays from Year 7 and 8 students was analyzed using five LLMs: GPT-3.5, GPT-4, o1, LLaMA 3-70B, and Mixtral 8x7B, aiming to provide in-depth insights into LLMs' scoring capabilities. Closed-source GPT models outperform open-source models in both internal consistency and alignment with human ratings, particularly excelling in language-related criteria. The novel o1 model outperforms all other LLMs, achieving Spearman's $r = .74$ with human assessments in the overall score, and an internal consistency of $ICC=.80$. These findings indicate that LLM-based assessment can be a useful tool to reduce teacher workload by supporting the evaluation of essays, especially with regard to language-related criteria. However, due to their tendency for higher scores, the models require further refinement to better capture aspects of content quality.
Motivation & Objective
- To assess the reliability and validity of large language models (LLMs) in scoring German student essays across multiple dimensions.
- To compare the performance of both open-source and closed-source LLMs against human teacher ratings on 10 predefined criteria.
- To evaluate the internal consistency and alignment with human assessments of LLMs in educational essay scoring.
- To identify strengths and limitations of LLMs in supporting teachers, particularly regarding language vs. content quality evaluation.
- To inform the development of AI-assisted essay scoring tools that reduce teacher workload while maintaining assessment quality.
Proposed method
- A corpus of 20 real-world German student essays from Year 7 and 8 was evaluated using five LLMs: GPT-3.5, GPT-4, o1, LLaMA 3-70B, and Mixtral 8x7B.
- Essays were scored by 37 human teachers on 10 predefined criteria, including plot logic, expression, and language accuracy.
- LLM evaluations were conducted without prompt engineering to ensure consistency and isolate model behavior.
- Statistical measures such as Spearman’s rank correlation (r) and Intraclass Correlation (ICC) were used to assess alignment and internal consistency.
- Comparative analysis focused on performance differences between closed-source (e.g., GPT-4, o1) and open-source (e.g., LLaMA 3, Mixtral) models.
- The study analyzed variance across multiple LLM runs to assess reliability and robustness of automated scoring.

Experimental results
Research questions
- RQ1How do closed-source and open-source LLMs compare in their alignment with human teacher ratings across 10 multidimensional criteria in German essay scoring?
- RQ2What is the internal consistency of LLM-generated scores, and how does it vary across different models?
- RQ3Which LLM performs best in replicating human teacher assessments, particularly in language-related vs. content-related criteria?
- RQ4To what extent do LLMs exhibit bias toward higher overall scores, and how does this affect their reliability in educational settings?
- RQ5How can LLMs be refined to better assess content quality while maintaining alignment with human evaluative standards?
Key findings
- The o1 model achieved the highest correlation with human ratings (Spearman’s r = .74) and the strongest internal consistency (ICC = .80), outperforming all other models.
- Closed-source models, especially GPT-4 and o1, significantly outperformed open-source models in both alignment with human ratings and internal consistency.
- Open-source models like LLaMA 3-70B and Mixtral 8x7B showed low variance but weak correlation with human ratings, indicating poor reliability for multidimensional assessment.
- LLMs demonstrated stronger performance on language-related criteria (e.g., expression, grammar) than on content-related criteria (e.g., plot logic, argumentation).
- The o1 model’s performance suggests that newer, reasoning-optimized LLMs can closely emulate human scoring, though they still tend to assign higher overall scores.
- The study highlights the need for further refinement in LLMs to improve content quality assessment and reduce scoring bias to ensure alignment with human standards.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.