[Paper Review] Automatic assessment of text-based responses in post-secondary education: A systematic review
This paper systematically reviews text-based automatic assessment systems in higher education, categorizing them into five IPO-based types and mapping educational focuses, motivations, and outcomes across 93 studies.
Text-based open-ended questions in academic formative and summative assessments help students become deep learners and prepare them to understand concepts for a subsequent conceptual assessment. However, grading text-based questions, especially in large courses, is tedious and time-consuming for instructors. Text processing models continue progressing with the rapid development of Artificial Intelligence (AI) tools and Natural Language Processing (NLP) algorithms. Especially after breakthroughs in Large Language Models (LLM), there is immense potential to automate rapid assessment and feedback of text-based responses in education. This systematic review adopts a scientific and reproducible literature search strategy based on the PRISMA process using explicit inclusion and exclusion criteria to study text-based automatic assessment systems in post-secondary education, screening 838 papers and synthesizing 93 studies. To understand how text-based automatic assessment systems have been developed and applied in education in recent years, three research questions are considered. All included studies are summarized and categorized according to a proposed comprehensive framework, including the input and output of the system, research motivation, and research outcomes, aiming to answer the research questions accordingly. Additionally, the typical studies of automated assessment systems, research methods, and application domains in these studies are investigated and summarized. This systematic review provides an overview of recent educational applications of text-based assessment systems for understanding the latest AI/NLP developments assisting in text-based assessments in higher education. Findings will particularly benefit researchers and educators incorporating LLMs such as ChatGPT into their educational activities.
Motivation & Objective
- Identify major automated text-based assessment system (TBAAS) types using an input-process-output framework.
- Characterize educational focuses, learning needs, and research motivations behind TBAAS.
- Synthesize reported outcomes, implications, and next steps for educational applications of AI/NLP in higher education.
Proposed method
- Adopt PRISMA-based literature search and screening across four databases (ACM DL, IEEE Xplore, Education Source, ASEE) from 2017–2023.
- Apply explicit inclusion/exclusion criteria to select primary empirical studies on text-based student answers in post-secondary settings.
- Extract data using an IPO (input-process-output) framework and open coding to identify themes and categorize TBAAS.
- Code and synthesize studies to map features, domains, motivations, and outcomes.
- Provide a comprehensive framework for understanding how AI/NLP advances support text-based assessment in higher education.
Experimental results
Research questions
- RQ1RQ1: What types of automated assessment systems can be identified using input, output, and processing framework?
- RQ2RQ2: What are the educational focuses and research motivations of studies with automated assessment systems?
- RQ3RQ3: What are the reported research outcomes in automated assessment systems, and what are the next steps for educational application?
Key findings
- Five TBAAS types were identified: Automatic Grading System (n=39), Automatic Classifier (n=22), Automatic Feedback System (n=20), Automated Writing Evaluation System (n=8), and Multimodal Evaluation System (n=4).
- Over half of the studies (55%) were in STEM domains, with computer science making up about 45% of STEM studies, 29% science, and 20% engineering; humanities accounted for ~32%, with English language studies most common (30%).
- Learning needs observed include: 1) improve assessment/grading/course evaluation (31 studies), 2) support specific content learning (25), 3) reduce grading time/effort (19), 4) support personalization/feedback (19).
- Research motivations commonly cited: automate grading/review/feedback (27 studies), analyze meaning in open-ended text (21), develop systems for learning (17), validate methods with human input (15), and test metrics (13).
- The review aggregates diverse methods (short answers, essays, constructed responses) and outputs (scores, labels, feedback, guidance), illustrating how NLP/AI techniques are applied to higher-education assessment.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.