Skip to main content
QUICK REVIEW

[논문 리뷰] GLAT: The Generative AI Literacy Assessment Test

Yueqiao Jin, Roberto Martínez‐Maldonado|arXiv (Cornell University)|2024. 11. 01.
Explainable Artificial Intelligence (XAI)인용 수 4
한 줄 요약

이 논문은 높은 교육 기관에서 생성형 AI(GenAI) 리터러시를 평가하기 위한 성과 기반 다중 선택 문제 20개로 구성된 GLAT를 소개한다. 고전적 시험 이론과 응답 이론을 기반으로 개발된 GLAT는 높은 신뢰성(Cronbach’s alpha = 0.80)과 타당성을 보이며, 실제 GenAI 지원 과제 수행 능력을 예측하는 데에 자가 보고 측정 방법보다 뛰어나다.

ABSTRACT

The rapid integration of generative artificial intelligence (GenAI) technology into education necessitates precise measurement of GenAI literacy to ensure that learners and educators possess the skills to engage with and critically evaluate this transformative technology effectively. Existing instruments often rely on self-reports, which may be biased. In this study, we present the GenAI Literacy Assessment Test (GLAT), a 20-item multiple-choice instrument developed following established procedures in psychological and educational measurement. Structural validity and reliability were confirmed with responses from 355 higher education students using classical test theory and item response theory, resulting in a reliable 2-parameter logistic (2PL) model (Cronbach's alpha = 0.80; omega total = 0.81) with a robust factor structure (RMSEA = 0.03; CFI = 0.97). Critically, GLAT scores were found to be significant predictors of learners' performance in GenAI-supported tasks, outperforming self-reported measures such as perceived ChatGPT proficiency and demonstrating external validity. These results suggest that GLAT offers a reliable and valid method for assessing GenAI literacy, with the potential to inform educational practices and policy decisions that aim to enhance learners' and educators' GenAI literacy, ultimately equipping them to navigate an AI-enhanced future.

연구 동기 및 목표

  • 높은 교육 기관에서 실제 GenAI 리터러시를 측정하는 데 신뢰성 있고 타당한 도구의 부족을 해결하기 위해.
  • 자기 보고식 AI 리터러시 설문 조사에 내재된 편향을 극복하기 위한 성과 기반 평가 도구를 개발하기 위해.
  • 고전적 시험 이론과 응답 이론을 포함한 철저한 심리측정 방법을 사용하여 도구를 검증하기 위해.
  • 실제 GenAI 지원 학습 과제에서의 성과와의 연계를 통해 외부 타당성을 확립하기 위해.
  • 교육자와 연구자들이 다양한 학술 맥락에서 GenAI 리터러시를 진단하고 향상시키기 위한 신뢰할 수 있는 도구를 제공하기 위해.

제안 방법

  • 기존 심리학 및 교육 측정 원리를 기반으로 한 20개의 다중 선택 문제로 구성된 도구를 개발하였다.
  • 심리측정 검증을 위한 데이터 수집을 위해 355명의 높은 교육 기관 학생들에게 테스트를 시행하였다.
  • 내적 일관성과 신뢰성을 평가하기 위해 고전적 시험 이론을 적용하여 Cronbach’s alpha = 0.80 및 omega total = 0.81을 산출하였다.
  • 항목 반응 이론을 사용하여 2파라미터 로지스틱(2PL) 모델을 적합시켰으며, RMSEA = 0.03 및 CFI = 0.97로 강력한 요인 구조를 확인하였다.
  • 채팅 봇과 시각적 분석을 포함한 맥락 특화 GenAI 과제에서의 성과와의 상관관계를 분석하여 외부 타당성을 검토하였다.
  • 기초 지식, 프롬프트 엔지니어링, GenAI 출력물의 윤리적 평가 등 다양한 인지 영역에서 도구의 유효성을 검증하였다.
Figure 1: The participant sample size and focus of each validation study.
Figure 1: The participant sample size and focus of each validation study.

실험 결과

연구 질문

  • RQ1성과 기반 평가가 높은 교육 기관 학생들의 실제 GenAI 리터러시를 신뢰성 있게 측정할 수 있는가?
  • RQ2자기 보고 측정 방법과 비교할 때 GLAT는 구조적 타당성과 내적 일관성 측면에서 어떻게 성과를 보이는가?
  • RQ3GLAT 점수는 실제 GenAI 지원 학습 과제에서 학습자의 성과를 어느 정도 예측할 수 있는가?
  • RQ4GLAT 점수는 실제 과제 성과를 예측하는 데에 자기 보고된 챗지피티 숙련도보다 어떻게 비교되는가?
  • RQ5현재의 자기 보고식 AI 리터러시 도구는 진정한 GenAI 능력을 포괄하는 데에 어떤 한계를 지니는가?

주요 결과

  • GLAT는 Cronbach’s alpha 0.80과 omega total 0.81을 보이며 높은 내적 일관성을 보이며, 높은 신뢰성을 나타냈다.
  • 도구는 루트 평균 제곱 오차의 근사치(RMSEA) 0.03과 비교적 적합 지수(CFI) 0.97로 뛰어난 구조적 타당성을 보였다.
  • GLAT 점수는 GenAI 지원 과제에서 학습자의 성과를 유의미하게 예측했으며, 자기 보고된 챗지피티 숙련도를 뛰어넘었다.
  • 2PL 모델은 데이터에 잘 적합되었으며, 테스트된 인구 집단 전반에서 도구의 심리측정적 탄탄함을 확인했다.
  • GLAT 점수와 실제 GenAI 맥락에서의 과제 성과 사이에 유의미한 정적 상관관계가 존재하여 외부 타당성이 입증되었다.
  • 본 연구는 성과 기반 평가 도구인 GLAT가 자기 보고에 비해 실제 GenAI 리터러시 능력을 더 정확하게 측정할 수 있음을 확인한다.
Figure 2: Visual analytics on teamwork in healthcare simulations, including: a) a bar chart of four prioritisation strategies, b) a social network diagram of communication behaviours among the actors, and c) a ward map showing individuals’ physical positions (hexagon), verbal communication duration
Figure 2: Visual analytics on teamwork in healthcare simulations, including: a) a bar chart of four prioritisation strategies, b) a social network diagram of communication behaviours among the actors, and c) a ward map showing individuals’ physical positions (hexagon), verbal communication duration

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.