Skip to main content
QUICK REVIEW

[论文解读] Automatic assessment of text-based responses in post-secondary education: A systematic review

Rujun Gao, Hillary Merzdorf|arXiv (Cornell University)|Aug 30, 2023
Artificial Intelligence in Healthcare and Education被引用 9
一句话总结

这篇论文系统性回顾高等教育中的基于文本的自动评估系统,将其归类为五种基于输入-过程-输出(IPO)的类型,并在93项研究中映射教育焦点、动机和结果。

ABSTRACT

Text-based open-ended questions in academic formative and summative assessments help students become deep learners and prepare them to understand concepts for a subsequent conceptual assessment. However, grading text-based questions, especially in large courses, is tedious and time-consuming for instructors. Text processing models continue progressing with the rapid development of Artificial Intelligence (AI) tools and Natural Language Processing (NLP) algorithms. Especially after breakthroughs in Large Language Models (LLM), there is immense potential to automate rapid assessment and feedback of text-based responses in education. This systematic review adopts a scientific and reproducible literature search strategy based on the PRISMA process using explicit inclusion and exclusion criteria to study text-based automatic assessment systems in post-secondary education, screening 838 papers and synthesizing 93 studies. To understand how text-based automatic assessment systems have been developed and applied in education in recent years, three research questions are considered. All included studies are summarized and categorized according to a proposed comprehensive framework, including the input and output of the system, research motivation, and research outcomes, aiming to answer the research questions accordingly. Additionally, the typical studies of automated assessment systems, research methods, and application domains in these studies are investigated and summarized. This systematic review provides an overview of recent educational applications of text-based assessment systems for understanding the latest AI/NLP developments assisting in text-based assessments in higher education. Findings will particularly benefit researchers and educators incorporating LLMs such as ChatGPT into their educational activities.

研究动机与目标

  • 使用输入-过程-输出框架识别主要的自动化文本评估系统(TBAAS)类型。
  • 描述 TBAAS 背后的教育焦点、学习需求和研究动机。
  • 综合报道的结果、含义,以及在高等教育中应用 AI/NLP 的下一步。

提出的方法

  • 采用基于 PRISMA 的文献检索与筛选,在四个数据库(ACM DL、IEEE Xplore、Education Source、ASEE)自2017-2023年期间。
  • 应用明确的纳入/排除标准,筛选关于后续教育环境中文本型学生回答的原始实证研究。
  • 使用 IPO(input-process-output)框架和开放编码提取数据,以识别主题并对 TBAAS 进行分类。
  • 对研究进行编码与综合,以绘制特征、领域、动机和结果的映射。
  • 提供一个全面的框架,帮助理解 AI/NLP 如何推动高等教育中的基于文本的评估。

实验结果

研究问题

  • RQ1RQ1:可以使用输入-输出-处理框架识别出哪些类型的自动评估系统?
  • RQ2RQ2:具有自动评估系统的研究在教育焦点和研究动机方面有哪些?
  • RQ3RQ3:在自动评估系统中报告的研究结果是什么,以及教育应用的下一步应该是什么?

主要发现

  • Five TBAAS types were identified: Automatic Grading System (n=39), Automatic Classifier (n=22), Automatic Feedback System (n=20), Automated Writing Evaluation System (n=8), and Multimodal Evaluation System (n=4).
  • Over half of the studies (55%) were in STEM domains, with computer science making up about 45% of STEM studies, 29% science, and 20% engineering; humanities accounted for ~32%, with English language studies most common (30%).
  • Learning needs observed include: 1) improve assessment/grading/course evaluation (31 studies), 2) support specific content learning (25), 3) reduce grading time/effort (19), 4) support personalization/feedback (19).
  • Research motivations commonly cited: automate grading/review/feedback (27 studies), analyze meaning in open-ended text (21), develop systems for learning (17), validate methods with human input (15), and test metrics (13).
  • The review aggregates diverse methods (short answers, essays, constructed responses) and outputs (scores, labels, feedback, guidance), illustrating how NLP/AI techniques are applied to higher-education assessment.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。