Skip to main content
QUICK REVIEW

[Paper Review] Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben

Rainer Muehlhoff, Marte Henningsen|arXiv (Cornell University)|Dec 9, 2024
Linguistic Education and Pedagogy4 citations
TL;DR

This study evaluates Fobizz's AI Grading Assistant, an LLM-based tool for automated homework assessment in German schools. Despite being marketed as objective and time-saving, the tool produces inconsistent, often nonsensical grades and feedback, with performance only improving when inputs are GPT-generated—highlighting fundamental flaws in LLM-based educational automation.

ABSTRACT

This study examines the AI-powered grading tool "AI Grading Assistant" by the German company Fobizz, designed to support teachers in evaluating and providing feedback on student assignments. Against the societal backdrop of an overburdened education system and rising expectations for artificial intelligence as a solution to these challenges, the investigation evaluates the tool's functional suitability through two test series. The results reveal significant shortcomings: The tool's numerical grades and qualitative feedback are often random and do not improve even when its suggestions are incorporated. The highest ratings are achievable only with texts generated by ChatGPT. False claims and nonsensical submissions frequently go undetected, while the implementation of some grading criteria is unreliable and opaque. Since these deficiencies stem from the inherent limitations of large language models (LLMs), fundamental improvements to this or similar tools are not immediately foreseeable. The study critiques the broader trend of adopting AI as a quick fix for systemic problems in education, concluding that Fobizz's marketing of the tool as an objective and time-saving solution is misleading and irresponsible. Finally, the study calls for systematic evaluation and subject-specific pedagogical scrutiny of the use of AI tools in educational contexts.

Motivation & Objective

  • To assess the functional suitability of Fobizz's AI Grading Assistant for automated homework evaluation in secondary education.
  • To investigate whether the tool reliably applies grading criteria and generates meaningful feedback.
  • To examine the impact of LLM limitations on the accuracy and fairness of automated assessment in educational contexts.
  • To critique the broader trend of adopting AI as a quick fix for systemic educational challenges.
  • To call for systematic, subject-specific evaluation of AI tools before institutional adoption in education.

Proposed method

  • Conducted two series of controlled tests using diverse student-written and GPT-generated assignments.
  • Evaluated the tool’s output against predefined grading rubrics and pedagogical standards.
  • Analyzed the consistency and coherence of numerical grades and qualitative feedback across multiple submissions.
  • Compared results from human-written student work with those from AI-generated texts to assess bias and reliability.
  • Assessed transparency and interpretability of the tool’s decision-making process.
  • Identified recurring errors such as false claims, logical inconsistencies, and misapplication of criteria.

Experimental results

Research questions

  • RQ1To what extent does Fobizz’s AI Grading Assistant produce reliable and consistent grades across varied student assignments?
  • RQ2How accurate and meaningful are the qualitative feedback comments generated by the tool?
  • RQ3Does the tool detect false or nonsensical content in student submissions, particularly when such content is present?
  • RQ4How does the tool’s performance vary between human-written and AI-generated student work?
  • RQ5To what degree is the tool’s decision-making process transparent and aligned with established pedagogical standards?

Key findings

  • The AI Grading Assistant frequently assigns grades and feedback that are inconsistent, random, or semantically incoherent.
  • Even after incorporating the tool’s suggestions, grading quality does not improve, indicating systemic flaws.
  • The highest possible grades are only attainable when inputs are generated by ChatGPT, suggesting a fundamental bias toward AI-generated text.
  • The tool fails to detect false claims and nonsensical content in student submissions, raising concerns about academic integrity.
  • Implementation of grading criteria is unreliable and opaque, with no clear explanation of how decisions are made.
  • These deficiencies stem from inherent limitations of large language models, making immediate improvements unlikely.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.