Skip to main content
QUICK REVIEW

[论文解读] Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben

Rainer Muehlhoff, Marte Henningsen|arXiv (Cornell University)|Dec 9, 2024
Linguistic Education and Pedagogy被引用 4
一句话总结

本研究评估了Fobizz的AI评分助手在德国中小学中的应用,该工具基于大语言模型,旨在实现作业的自动化评分。尽管该工具被宣传为客观且节省时间,但其实际表现却存在评分不一致、反馈常无意义等问题,且仅在输入为GPT生成内容时性能才有所提升,凸显出基于大语言模型的教育自动化存在根本性缺陷。

ABSTRACT

This study examines the AI-powered grading tool "AI Grading Assistant" by the German company Fobizz, designed to support teachers in evaluating and providing feedback on student assignments. Against the societal backdrop of an overburdened education system and rising expectations for artificial intelligence as a solution to these challenges, the investigation evaluates the tool's functional suitability through two test series. The results reveal significant shortcomings: The tool's numerical grades and qualitative feedback are often random and do not improve even when its suggestions are incorporated. The highest ratings are achievable only with texts generated by ChatGPT. False claims and nonsensical submissions frequently go undetected, while the implementation of some grading criteria is unreliable and opaque. Since these deficiencies stem from the inherent limitations of large language models (LLMs), fundamental improvements to this or similar tools are not immediately foreseeable. The study critiques the broader trend of adopting AI as a quick fix for systemic problems in education, concluding that Fobizz's marketing of the tool as an objective and time-saving solution is misleading and irresponsible. Finally, the study calls for systematic evaluation and subject-specific pedagogical scrutiny of the use of AI tools in educational contexts.

研究动机与目标

  • 评估Fobizz的AI评分助手在中学教育中用于自动化作业评估的功能适用性。
  • 调查该工具是否能可靠地应用评分标准并生成有意义的反馈。
  • 探讨大语言模型的局限性对教育场景中自动化评估准确性与公平性的影响。
  • 批判将AI视为解决系统性教育挑战的快速解决方案这一普遍趋势。
  • 呼吁在教育机构正式采用AI工具前,开展系统性、学科特定的评估。

提出的方法

  • 通过多样化的学生原创作业与GPT生成的作业,开展两轮受控测试。
  • 依据预设的评分量规和教学标准,评估该工具的输出结果。
  • 分析在多份作业提交中,数值评分与定性反馈的一致性与连贯性。
  • 将人工撰写的作业结果与AI生成文本的结果进行对比,以评估偏差与可靠性。
  • 评估该工具决策过程的透明度与可解释性。
  • 识别出重复出现的错误,如错误陈述、逻辑矛盾以及评分标准的误用。

实验结果

研究问题

  • RQ1Fobizz的AI评分助手在不同学生作业中,其评分的可靠性与一致性如何?
  • RQ2该工具生成的定性反馈意见在准确性和意义性方面表现如何?
  • RQ3当学生作业中存在虚假或无意义内容时,该工具能否有效识别?
  • RQ4该工具在处理人工撰写与AI生成的作业时,表现有何差异?
  • RQ5该工具的决策过程在多大程度上具备透明度,并符合既定的教学标准?

主要发现

  • AI评分助手频繁给出不一致、随机或语义上不连贯的评分与反馈。
  • 即使采纳了该工具的建议,评分质量也未见提升,表明其存在系统性缺陷。
  • 只有当输入为ChatGPT生成时,才能获得最高分,表明该工具对AI生成文本存在根本性偏见。
  • 该工具未能识别学生作业中的虚假陈述与无意义内容,引发对学术诚信的担忧。
  • 评分标准的实施不可靠且不透明,缺乏对决策依据的清晰说明。
  • 这些缺陷源于大语言模型的固有局限性,短期内难以实现根本性改善。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。