Skip to main content
QUICK REVIEW

[Paper Review] Towards LLM-based Autograding for Short Textual Answers

Johannes Schneider, Bernd Schenk|arXiv (Cornell University)|Sep 9, 2023
Artificial Intelligence in Healthcare and EducationMedicine3 citations
TL;DR

This paper evaluates large language models (LLMs) for automated short-text answer grading across multiple languages and courses, demonstrating that while LLMs offer a valuable complementary perspective, they are not yet ready for fully autonomous grading without human oversight due to reliability and bias concerns. The study highlights LLMs' potential to assist educators in validating grading procedures.

ABSTRACT

Grading exams is an important, labor-intensive, subjective, repetitive, and frequently challenging task. The feasibility of autograding textual responses has greatly increased thanks to the availability of large language models (LLMs) such as ChatGPT and the substantial influx of data brought about by digitalization. However, entrusting AI models with decision-making roles raises ethical considerations, mainly stemming from potential biases and issues related to generating false information. Thus, in this manuscript, we provide an evaluation of a large language model for the purpose of autograding, while also highlighting how LLMs can support educators in validating their grading procedures. Our evaluation is targeted towards automatic short textual answers grading (ASAG), spanning various languages and examinations from two distinct courses. Our findings suggest that while "out-of-the-box" LLMs provide a valuable tool to provide a complementary perspective, their readiness for independent automated grading remains a work in progress, necessitating human oversight.

Motivation & Objective

  • To assess the feasibility and reliability of LLMs for grading short textual answers in educational settings.
  • To evaluate LLM performance across diverse languages and examination types in two distinct academic courses.
  • To explore how LLMs can support educators in validating and auditing their own grading decisions.
  • To identify limitations and ethical risks—particularly bias and hallucination—when deploying LLMs in automated grading.
  • To provide evidence-based recommendations for integrating LLMs into grading workflows with appropriate human oversight.

Proposed method

  • The authors applied a pre-trained LLM to grade short-answer responses from two university-level courses, covering multiple languages.
  • Grading was performed using prompt engineering to elicit consistent, rubric-aligned evaluations from the LLM.
  • The LLM's scores were compared against human-graded benchmarks to assess accuracy and consistency.
  • The evaluation included both quantitative metrics (e.g., correlation with human scores) and qualitative analysis of LLM reasoning.
  • The study used a multi-stage process: prompt design, LLM inference, score aggregation, and comparison with human-graded gold standards.
  • Ethical considerations such as bias and hallucination were analyzed through case studies of LLM-generated feedback.

Experimental results

Research questions

  • RQ1To what extent can LLMs produce grading scores that correlate with human-graded benchmarks for short textual answers?
  • RQ2How does LLM performance vary across different languages and subject domains in educational assessments?
  • RQ3In what ways can LLMs assist educators in validating or auditing their own grading decisions?
  • RQ4What are the primary risks—such as bias or hallucination—when using LLMs for automated grading?
  • RQ5How can LLMs be integrated into grading workflows to support, rather than replace, human graders?

Key findings

  • LLMs demonstrated moderate to high correlation with human-graded scores, indicating potential as a complementary grading tool.
  • Performance varied significantly across languages, with lower accuracy in non-English and less-resourced languages.
  • The LLM occasionally generated plausible but incorrect feedback, indicating risks of hallucination and bias.
  • Human oversight was consistently necessary to correct errors and ensure fairness and consistency in grading outcomes.
  • LLMs provided useful feedback that helped educators identify inconsistencies in their own grading patterns.
  • The study found that 'out-of-the-box' LLMs are not yet reliable for independent automated grading without fine-tuning or rigorous validation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.