Skip to main content
QUICK REVIEW

[Paper Review] GLAT: The Generative AI Literacy Assessment Test

Yueqiao Jin, Roberto Martínez‐Maldonado|arXiv (Cornell University)|Nov 1, 2024
Explainable Artificial Intelligence (XAI)4 citations
TL;DR

This paper introduces GLAT, a 20-item performance-based multiple-choice test to assess generative AI (GenAI) literacy in higher education. Developed using classical test theory and item response theory, GLAT demonstrates strong reliability (Cronbach’s alpha = 0.80) and validity, outperforming self-reported measures in predicting actual performance in GenAI-supported tasks.

ABSTRACT

The rapid integration of generative artificial intelligence (GenAI) technology into education necessitates precise measurement of GenAI literacy to ensure that learners and educators possess the skills to engage with and critically evaluate this transformative technology effectively. Existing instruments often rely on self-reports, which may be biased. In this study, we present the GenAI Literacy Assessment Test (GLAT), a 20-item multiple-choice instrument developed following established procedures in psychological and educational measurement. Structural validity and reliability were confirmed with responses from 355 higher education students using classical test theory and item response theory, resulting in a reliable 2-parameter logistic (2PL) model (Cronbach's alpha = 0.80; omega total = 0.81) with a robust factor structure (RMSEA = 0.03; CFI = 0.97). Critically, GLAT scores were found to be significant predictors of learners' performance in GenAI-supported tasks, outperforming self-reported measures such as perceived ChatGPT proficiency and demonstrating external validity. These results suggest that GLAT offers a reliable and valid method for assessing GenAI literacy, with the potential to inform educational practices and policy decisions that aim to enhance learners' and educators' GenAI literacy, ultimately equipping them to navigate an AI-enhanced future.

Motivation & Objective

  • To address the lack of reliable, valid instruments for measuring actual GenAI literacy in higher education.
  • To develop a performance-based assessment tool that overcomes the biases inherent in self-reported AI literacy surveys.
  • To validate the instrument using rigorous psychometric methods, including classical test theory and item response theory.
  • To establish external validity by linking GLAT scores to performance in real GenAI-supported learning tasks.
  • To provide educators and researchers with a reliable tool to diagnose and enhance GenAI literacy across academic contexts.

Proposed method

  • Developed a 20-item multiple-choice instrument based on established psychological and educational measurement principles.
  • Administered the test to 355 higher education students to collect data for psychometric validation.
  • Applied classical test theory to assess internal consistency and reliability, yielding Cronbach’s alpha = 0.80 and omega total = 0.81.
  • Used item response theory to fit a 2-parameter logistic (2PL) model, confirming a robust factor structure with RMSEA = 0.03 and CFI = 0.97.
  • Conducted external validity analysis by correlating GLAT scores with performance in context-specific GenAI tasks involving chatbots and visual analytics.
  • Validated the instrument across diverse cognitive domains, including foundational knowledge, prompt engineering, and ethical evaluation of GenAI outputs.
Figure 1: The participant sample size and focus of each validation study.
Figure 1: The participant sample size and focus of each validation study.

Experimental results

Research questions

  • RQ1Can a performance-based assessment reliably measure actual GenAI literacy in higher education students?
  • RQ2How does GLAT perform in terms of structural validity and internal consistency compared to self-reported measures?
  • RQ3To what extent do GLAT scores predict learners’ performance in real GenAI-supported learning tasks?
  • RQ4How does GLAT compare to self-reported proficiency in predicting actual task performance?
  • RQ5What are the limitations of current self-reported AI literacy instruments in capturing true GenAI competencies?

Key findings

  • GLAT demonstrated strong internal consistency with Cronbach’s alpha of 0.80 and omega total of 0.81, indicating high reliability.
  • The instrument exhibited excellent structural validity, with a root mean square error of approximation (RMSEA) of 0.03 and comparative fit index (CFI) of 0.97.
  • GLAT scores were significant predictors of learners’ performance in GenAI-supported tasks, outperforming self-reported ChatGPT proficiency.
  • The 2PL model fit the data well, confirming the instrument’s psychometric robustness across the tested population.
  • External validity was established through a significant positive relationship between GLAT scores and actual task performance in GenAI contexts.
  • The study confirms that performance-based assessments like GLAT are more accurate than self-reports in measuring real GenAI literacy competencies.
Figure 2: Visual analytics on teamwork in healthcare simulations, including: a) a bar chart of four prioritisation strategies, b) a social network diagram of communication behaviours among the actors, and c) a ward map showing individuals’ physical positions (hexagon), verbal communication duration
Figure 2: Visual analytics on teamwork in healthcare simulations, including: a) a bar chart of four prioritisation strategies, b) a social network diagram of communication behaviours among the actors, and c) a ward map showing individuals’ physical positions (hexagon), verbal communication duration

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.