Skip to main content
QUICK REVIEW

[论文解读] AI and Machine Learning for Next Generation Science Assessments

Xiaoming Zhai|arXiv (Cornell University)|Apr 23, 2024
Genetics, Bioinformatics, and Biomedical Research被引用 6
一句话总结

本章评估AI/ML如何使下一代三维对齐的科学评估成为可能,提出用于评分准确性的框架,并讨论未来方向与挑战。

ABSTRACT

This chapter focuses on the transformative role of Artificial Intelligence (AI) and Machine Learning (ML) in science assessments. The paper begins with a discussion of the Framework for K-12 Science Education, which calls for a shift from conceptual learning to knowledge-in-use. This shift necessitates the development of new types of assessments that align with the Framework's three dimensions: science and engineering practices, disciplinary core ideas, and crosscutting concepts. The paper further highlights the limitations of traditional assessment methods like multiple-choice questions, which often fail to capture the complexities of scientific thinking and three-dimensional learning in science. It emphasizes the need for performance-based assessments that require students to engage in scientific practices like modeling, explanation, and argumentation. The paper achieves three major goals: reviewing the current state of ML-based assessments in science education, introducing a framework for scoring accuracy in ML-based automatic assessments, and discussing future directions and challenges. It delves into the evolution of ML-based automatic scoring systems, discussing various types of ML, like supervised, unsupervised, and semi-supervised learning. These systems can provide timely and objective feedback, thus alleviating the burden on teachers. The paper concludes by exploring pre-trained models like BERT and finetuned ChatGPT, which have shown promise in assessing students' written responses effectively.

研究动机与目标

  • 评估科学教育中基于ML的评估的现状与机会。
  • 提出一个在ML自动评估评分中考虑准确性的框架。
  • 识别在科学评估中部署ML的挑战、方向和伦理考量。

提出的方法

  • 回顾科学评估中基于ML的自动评分系统的演变。
  • 描述Machine-Human Agreement (MHA)框架及其在五个类别中的调节因素。
  • 讨论自动评分中的预训练模型(例如BERT)以及零-shot/少-shot方法。
  • 总结监督、无监督、半监督和零-shot学习在评分中的利与弊。
  • 强调有效性、公平性等问题,以及采用与持续评估的指南。
Figure 1: Assessment task: “Red dye diffusion” item screenshot (left), response interface (right), and a student response (bottom)(adapted from (Zhai, He, & Krajcik, 2022))
Figure 1: Assessment task: “Red dye diffusion” item screenshot (left), response interface (right), and a student response (bottom)(adapted from (Zhai, He, & Krajcik, 2022))

实验结果

研究问题

  • RQ1科学教育中基于ML的评估的当前状态与演变是什么?
  • RQ2如何为ML基于评估的评分准确性(MHA)框架化并改进?
  • RQ3预训练模型和零-shot/少-shot方法在科学任务的自动评分中起到什么作用?
  • RQ4对于ML基于的下一代科学评估,主要挑战、伦理考虑和未来方向是什么?

主要发现

  • 基于ML的评估可以提供及时且客观的反馈,减轻教师工作量。
  • 五类别框架(外部特征、内部特征、考生特征、训练/验证方法、技术特征)调节机器-人类一致性(MHA)。
  • 像BERT这样的预训练模型和微调的ChatGPT在科学教育中的书面回答评分方面显示出潜力。
  • 零-shot和少-shot方法可以达到非平凡的评分准确性(例如 MeNSP Cohen’s Kappa 0.30–0.57;少-shot 0.38)。
  • 微调的领域特定GPT-3.5模型在多项任务上的表现优于BERT,报告的平均准确率提升(如在多任务中的9.1%)。
  • 主要挑战包括模型泛化、数据不平衡,以及在基于ML的评分中对用户指南和透明度的需求。
Figure 2: Automatic scoring accuracy for science assessments involving different scientific practices
Figure 2: Automatic scoring accuracy for science assessments involving different scientific practices

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。