[論文レビュー] Automatic assessment of text-based responses in post-secondary education: A systematic review
この論文は、93件の研究にわたって高等教育におけるテキストベースの自動評価システムを体系的にレビューし、教育的焦点、動機、成果を5つの IPOベースのタイプに分類してマッピングします。
Text-based open-ended questions in academic formative and summative assessments help students become deep learners and prepare them to understand concepts for a subsequent conceptual assessment. However, grading text-based questions, especially in large courses, is tedious and time-consuming for instructors. Text processing models continue progressing with the rapid development of Artificial Intelligence (AI) tools and Natural Language Processing (NLP) algorithms. Especially after breakthroughs in Large Language Models (LLM), there is immense potential to automate rapid assessment and feedback of text-based responses in education. This systematic review adopts a scientific and reproducible literature search strategy based on the PRISMA process using explicit inclusion and exclusion criteria to study text-based automatic assessment systems in post-secondary education, screening 838 papers and synthesizing 93 studies. To understand how text-based automatic assessment systems have been developed and applied in education in recent years, three research questions are considered. All included studies are summarized and categorized according to a proposed comprehensive framework, including the input and output of the system, research motivation, and research outcomes, aiming to answer the research questions accordingly. Additionally, the typical studies of automated assessment systems, research methods, and application domains in these studies are investigated and summarized. This systematic review provides an overview of recent educational applications of text-based assessment systems for understanding the latest AI/NLP developments assisting in text-based assessments in higher education. Findings will particularly benefit researchers and educators incorporating LLMs such as ChatGPT into their educational activities.
研究の動機と目的
- 入力-処理-出力フレームワークを用いて主要な自動テキストベースの評価システム(TBAAS)のタイプを特定する。
- TBAASの背後にある教育的焦点、学習ニーズ、および研究動機を特徴づける。
- 高等教育におけるAI/NLPの教育応用に関する報告された成果、意味、および今後の方針を統合する。
提案手法
- 2017–2023年を対象に、4つのデータベース(ACM DL、IEEE Xplore、Education Source、ASEE)を用いたPRISMAベースの文献検索とスクリーニングを採用する。
- ポストセカンダリ設定におけるテキストベースの学生回答に関する一次実証研究を選択するため、明確な包含/排除基準を適用する。
- IPO(input-process-output)フレームワークとオープンコーディングを用いてデータを抽出し、テーマを特定しTBAASを分類する。
- 研究をコード化し統合して特徴、ドメイン、動機、成果をマッピングする。
- 高等教育におけるテキストベースの評価を支えるAI/NLPの進展を理解するための包括的なフレームワークを提供する。
実験結果
リサーチクエスチョン
- RQ1RQ1: 入力-出力-処理フレームワークを用いて識別できる自動評価システムのタイプは何ですか?
- RQ2RQ2: 自動評価システムを用いた研究の教育的焦点と研究動機は何ですか?
- RQ3RQ3: 自動評価システムにおける報告された研究成果は何であり、教育への適用のための今後のステップは何ですか?
主な発見
- 5つの TBAAS タイプが特定されました:Automatic Grading System (n=39)、Automatic Classifier (n=22)、Automatic Feedback System (n=20)、Automated Writing Evaluation System (n=8)、Multimodal Evaluation System (n=4)。
- 研究の過半数(55%)はSTEM分野に属し、コンピュータサイエンスがSTEM研究のおよそ45%、29%が科学、20%が工学を占める;人文学は約32%、英語言語学習に関する研究が最も多く(30%)を占めた。
- 学習ニーズには、1) 評定/評価/コース評価の改善(31件の研究)、2) 特定の内容学習の支援(25件)、3) 評定時間/労力の削減(19件)、4) パーソナライゼーション/フィードバックの支援(19件)が含まれます。
- 研究動機として一般的に挙げられるのは:自動採点/レビュー/フィードバックを自動化(27件の研究)、オープンエンドテキストの意味を分析(21)、学習のためのシステムを開発(17)、人間の入力で方法を検証(15)、指標をテスト(13)です。
- 本レビューは短答、エッセイ、構成化回答など多様な方法と、スコア、ラベル、フィードバック、指針といったアウトプットを統合しており、NLP/AI技術が高等教育の評価にどのように適用されているかを示しています。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。