Skip to main content
QUICK REVIEW

[Paper Review] AI and Machine Learning for Next Generation Science Assessments

Xiaoming Zhai|arXiv (Cornell University)|Apr 23, 2024
Genetics, Bioinformatics, and Biomedical Research6 citations
TL;DR

The chapter reviews how AI/ML can enable next-generation, three-dimensionally aligned science assessments, proposes a framework for scoring accuracy, and discusses future directions and challenges.

ABSTRACT

This chapter focuses on the transformative role of Artificial Intelligence (AI) and Machine Learning (ML) in science assessments. The paper begins with a discussion of the Framework for K-12 Science Education, which calls for a shift from conceptual learning to knowledge-in-use. This shift necessitates the development of new types of assessments that align with the Framework's three dimensions: science and engineering practices, disciplinary core ideas, and crosscutting concepts. The paper further highlights the limitations of traditional assessment methods like multiple-choice questions, which often fail to capture the complexities of scientific thinking and three-dimensional learning in science. It emphasizes the need for performance-based assessments that require students to engage in scientific practices like modeling, explanation, and argumentation. The paper achieves three major goals: reviewing the current state of ML-based assessments in science education, introducing a framework for scoring accuracy in ML-based automatic assessments, and discussing future directions and challenges. It delves into the evolution of ML-based automatic scoring systems, discussing various types of ML, like supervised, unsupervised, and semi-supervised learning. These systems can provide timely and objective feedback, thus alleviating the burden on teachers. The paper concludes by exploring pre-trained models like BERT and finetuned ChatGPT, which have shown promise in assessing students' written responses effectively.

Motivation & Objective

  • Assess current state and opportunities of ML-based assessments in science education.
  • Propose a framework to account for scoring accuracy in ML-based automatic assessments.
  • Identify challenges, directions, and ethical considerations for deploying ML in science assessment.

Proposed method

  • Review the evolution of ML-based automatic scoring systems for science assessments.
  • Describe a framework for Machine-Human Agreement (MHA) and factors moderating it across five categories.
  • Discuss pre-trained models (e.g., BERT) and zero-shot/few-shot approaches in automatic scoring.
  • Summarize benefits and limitations of supervised, unsupervised, semi-supervised, and zero-shot learning in scoring.
  • Highlight issues of validity, fairness, and guidelines for adoption and ongoing evaluation.
Figure 1: Assessment task: “Red dye diffusion” item screenshot (left), response interface (right), and a student response (bottom)(adapted from (Zhai, He, & Krajcik, 2022))
Figure 1: Assessment task: “Red dye diffusion” item screenshot (left), response interface (right), and a student response (bottom)(adapted from (Zhai, He, & Krajcik, 2022))

Experimental results

Research questions

  • RQ1What is the current state and evolution of ML-based assessments in science education?
  • RQ2How can scoring accuracy (MHA) be framed and improved for ML-based assessments?
  • RQ3What role do pre-trained models and zero-shot/few-shot approaches play in automatic scoring of science tasks?
  • RQ4What are the major challenges, ethical considerations, and future directions for ML-based Next Generation Science Assessments?

Key findings

  • ML-based assessments can provide timely and objective feedback, reducing teacher workload.
  • A five-category framework (external features, internal features, examinee features, training/validation approaches, technical features) modulates machine-human agreement (MHA).
  • Pre-trained models like BERT and fine-tuned ChatGPT show promise for scoring written responses in science education.
  • Zero-shot and few-shot approaches can achieve non-trivial scoring accuracy (e.g., MeNSP Cohen’s Kappa 0.30–0.57; few-shot 0.38).
  • Fine-tuned domain-specific GPT-3.5 models can outperform BERT on multiple tasks, with reported average accuracy gains (e.g., 9.1% across tasks).
  • Key challenges include model generalizability, unbalanced data, and need for user guidelines and transparency in ML-based scoring.
Figure 2: Automatic scoring accuracy for science assessments involving different scientific practices
Figure 2: Automatic scoring accuracy for science assessments involving different scientific practices

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.