[论文解读] GLAT: The Generative AI Literacy Assessment Test
本文介绍了GLAT,一种用于评估高等教育机构生成式人工智能(GenAI)素养的20道题表现性多选题测试。该测试基于经典测试理论和项目反应理论开发,表现出较强的信度(Cronbach’s alpha = 0.80)和效度,其预测实际GenAI支持任务表现的能力优于自报测量方法。
The rapid integration of generative artificial intelligence (GenAI) technology into education necessitates precise measurement of GenAI literacy to ensure that learners and educators possess the skills to engage with and critically evaluate this transformative technology effectively. Existing instruments often rely on self-reports, which may be biased. In this study, we present the GenAI Literacy Assessment Test (GLAT), a 20-item multiple-choice instrument developed following established procedures in psychological and educational measurement. Structural validity and reliability were confirmed with responses from 355 higher education students using classical test theory and item response theory, resulting in a reliable 2-parameter logistic (2PL) model (Cronbach's alpha = 0.80; omega total = 0.81) with a robust factor structure (RMSEA = 0.03; CFI = 0.97). Critically, GLAT scores were found to be significant predictors of learners' performance in GenAI-supported tasks, outperforming self-reported measures such as perceived ChatGPT proficiency and demonstrating external validity. These results suggest that GLAT offers a reliable and valid method for assessing GenAI literacy, with the potential to inform educational practices and policy decisions that aim to enhance learners' and educators' GenAI literacy, ultimately equipping them to navigate an AI-enhanced future.
研究动机与目标
- 为解决高等教育中缺乏可靠、有效的实际GenAI素养测量工具的问题。
- 开发一种基于表现的评估工具,以克服自报人工智能素养调查中固有的偏差。
- 使用严谨的心理测量方法(包括经典测试理论和项目反应理论)对工具进行验证。
- 通过将GLAT分数与真实GenAI支持学习任务中的表现关联,建立外部效度。
- 为教育工作者和研究人员提供一种可靠的工具,用于诊断和提升各类学术情境下的GenAI素养。
提出的方法
- 基于既定的心理学与教育测量原理,开发了一种20道题的多选题量表。
- 向355名高等教育学生发放测试,以收集用于心理测量验证的数据。
- 应用经典测试理论评估内部一致性与信度,得出Cronbach’s alpha = 0.80,omega total = 0.81。
- 使用项目反应理论拟合两参数逻辑模型(2PL),确认了具有RMSEA = 0.03和CFI = 0.97的稳健因子结构。
- 通过将GLAT分数与涉及聊天机器人和视觉分析的特定情境GenAI任务表现进行相关分析,开展外部效度分析。
- 在多种认知领域对工具进行了验证,包括基础知识、提示工程以及对GenAI输出的伦理评估。

实验结果
研究问题
- RQ1表现性评估能否可靠地测量高等教育学生的真实GenAI素养?
- RQ2与自报测量相比,GLAT在结构效度和内部一致性方面表现如何?
- RQ3GLAT分数在多大程度上能预测学习者在真实GenAI支持学习任务中的表现?
- RQ4GLAT在预测实际任务表现方面,与自报熟练度相比表现如何?
- RQ5当前自报人工智能素养工具在捕捉真实GenAI能力方面存在哪些局限?
主要发现
- GLAT表现出良好的内部一致性,Cronbach’s alpha为0.80,omega total为0.81,表明其信度较高。
- 该工具表现出优异的结构效度,近似误差均方根(RMSEA)为0.03,比较拟合指数(CFI)为0.97。
- GLAT分数是学习者在GenAI支持任务中表现的显著预测指标,优于自报的ChatGPT熟练度。
- 2PL模型对数据的拟合良好,证实了该工具在受测人群中的心理测量稳健性。
- 通过GLAT分数与真实GenAI情境下任务表现之间的显著正相关关系,确立了外部效度。
- 本研究证实,像GLAT这样的表现性评估比自报方法更能准确衡量真实的GenAI素养能力。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。