[논문 리뷰] Automatic assessment of text-based responses in post-secondary education: A systematic review
이 논문은 고등교육에서 텍스트 기반 자동 평가 시스템을 체계적으로 검토하여 다섯 IPO 기반 유형으로 분류하고 93개 연구 전반에 걸친 교육적 초점, 동기 및 결과를 매핑합니다.
Text-based open-ended questions in academic formative and summative assessments help students become deep learners and prepare them to understand concepts for a subsequent conceptual assessment. However, grading text-based questions, especially in large courses, is tedious and time-consuming for instructors. Text processing models continue progressing with the rapid development of Artificial Intelligence (AI) tools and Natural Language Processing (NLP) algorithms. Especially after breakthroughs in Large Language Models (LLM), there is immense potential to automate rapid assessment and feedback of text-based responses in education. This systematic review adopts a scientific and reproducible literature search strategy based on the PRISMA process using explicit inclusion and exclusion criteria to study text-based automatic assessment systems in post-secondary education, screening 838 papers and synthesizing 93 studies. To understand how text-based automatic assessment systems have been developed and applied in education in recent years, three research questions are considered. All included studies are summarized and categorized according to a proposed comprehensive framework, including the input and output of the system, research motivation, and research outcomes, aiming to answer the research questions accordingly. Additionally, the typical studies of automated assessment systems, research methods, and application domains in these studies are investigated and summarized. This systematic review provides an overview of recent educational applications of text-based assessment systems for understanding the latest AI/NLP developments assisting in text-based assessments in higher education. Findings will particularly benefit researchers and educators incorporating LLMs such as ChatGPT into their educational activities.
연구 동기 및 목표
- 입력-프로세스-출력 프레임워크를 사용하여 주요 자동화된 텍스트 기반 평가 시스템(TBAAS) 유형을 식별한다.
- TBAAS의 교육적 초점, 학습 필요성 및 연구 동기를 특징짓다.
- 고등교육에서 AI/NLP의 교육 적용에 대한 보고된 결과, 시사점 및 향후 단계를 합성한다.
제안 방법
- PRISMA 기반 문헌 검색 및 스크리닝을 2017–2023년 동안 네 가지 데이터베이스(ACM DL, IEEE Xplore, Education Source, ASEE)에 대해 채택한다.
- 포스트세컨더리 설정의 텍스트 기반 학생 답변에 대한 일차 경험적 연구를 선택하기 위해 명시적 포함/배제 기준을 적용한다.
- IPO (input-process-output) 프레임워크와 개방 코딩으로 데이터를 추출하여 주제와 범주화를 통해 TBAAS를 식별한다.
- 연구를 코드화하고 합성하여 특징, 도메인, 동기 및 결과를 매핑한다.
- 고등교육에서 AI/NLP 발전이 텍스트 기반 평가를 어떻게 지원하는지 이해하기 위한 포괄적 프레임워크를 제공한다.
실험 결과
연구 질문
- RQ1RQ1: 입력-처리 프레임워크를 사용하여 자동 평가 시스템의 어떤 유형을 식별할 수 있는가?
- RQ2RQ2: 자동 평가 시스템을 가진 연구의 교육적 초점과 연구 동기는 무엇인가?
- RQ3RQ3: 자동 평가 시스템에서 보고된 연구 결과는 무엇이며 교육적 적용의 다음 단계는 무엇인가?
주요 결과
- 다섯 가지 TBAAS 유형이 식별되었습니다: Automatic Grading System (n=39), Automatic Classifier (n=22), Automatic Feedback System (n=20), Automated Writing Evaluation System (n=8), 그리고 Multimodal Evaluation System (n=4).
- 연구의 절반 이상(55%)은 STEM 도메인에 속했고, 컴퓨터 과학은 STEM 연구의 약 45%, 29%는 과학, 20%는 공학으로 구성되며; 인문학은 약 32%를 차지했고 영어 언어 연구가 가장 흔했다(30%).
- 학습 필요성으로 관찰된 것: 1) 평가/채점/강의 평가 개선 (31 연구), 2) 특정 내용 학습 지원 (25), 3) 채점 시간/노력 감소 (19), 4) 개인화/피드백 지원 (19).
- 연구 동기로 자주 인용된 것: 자동 채점/리뷰/피드백 (27 연구), 개방형 텍스트의 의미 분석 (21), 학습 시스템 개발 (17), 인간 입력으로 방법 검증 (15), 지표 테스트 (13).
- 리뷰는 다양한 방법(짧은 답변, 에세이, 구성된 답변)과 산출물(점수, 라벨, 피드백, 가이드)을 집계하며 NLP/AI 기술이 고등교육 평가에 어떻게 적용되는지 보여준다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.