Skip to main content
QUICK REVIEW

[Paper Review] Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions

Hamdireza Rouzegar, Masoud Makrehchi|arXiv (Cornell University)|Jun 20, 2024
Intelligent Tutoring Systems and Adaptive Learning4 citations
TL;DR

This study investigates how GPT-3.5 and GPT-4 can generate customized, adaptive math test questions for Grade 9 students using an active learning framework. By positioning GPT-4 as a 'teacher' and GPT-3.5 as a 'student' in an iterative feedback loop, the models simulate real-time learning adaptation, with GPT-4 showing superior question quality and GPT-3.5 demonstrating improved performance on complex problems after instruction, highlighting the potential of LLMs to enhance personalized education.

ABSTRACT

This study investigates how LLMs, specifically GPT-3.5 and GPT-4, can develop tailored questions for Grade 9 math, aligning with active learning principles. By utilizing an iterative method, these models adjust questions based on difficulty and content, responding to feedback from a simulated 'student' model. A novel aspect of the research involved using GPT-4 as a 'teacher' to create complex questions, with GPT-3.5 as the 'student' responding to these challenges. This setup mirrors active learning, promoting deeper engagement. The findings demonstrate GPT-4's superior ability to generate precise, challenging questions and notable improvements in GPT-3.5's ability to handle more complex problems after receiving instruction from GPT-4. These results underscore the potential of LLMs to mimic and enhance active learning scenarios, offering a promising path for AI in customized education. This research contributes to understanding how AI can support personalized learning experiences, highlighting the need for further exploration in various educational contexts

Motivation & Objective

  • To evaluate the effectiveness of GPT-3.5 and GPT-4 in generating personalized, curriculum-aligned test questions for Grade 9 mathematics.
  • To simulate an active learning environment by using GPT-4 as a 'teacher' and GPT-3.5 as a 'student' to assess iterative learning and adaptation.
  • To examine how LLMs respond to increasing question difficulty and feedback, emulating human-like learning progression.
  • To assess the potential of large language models to support personalized education through dynamic, adaptive question generation.
  • To contribute insights into the integration of LLMs in educational technology for scalable, student-centered learning experiences.

Proposed method

  • An iterative question-generation and feedback loop was implemented, where GPT-3.5 responded to questions crafted by GPT-4, with adjustments based on response quality and difficulty.
  • GPT-4 was used as a 'teacher' to generate increasingly complex questions in 'Numbers' and 'Financial Mathematics' for Grade 9, targeting specific cognitive levels.
  • GPT-3.5 acted as a 'student' responding to these questions, with its answers evaluated for correctness and depth to inform subsequent question refinement.
  • The process mimicked active learning by adjusting question difficulty and content based on the model’s performance, simulating formative assessment.
  • A simulated 'student' model provided feedback, enabling the system to iteratively improve question quality and alignment with learning objectives.
  • Evaluation focused on accuracy, complexity, and adaptability of questions generated across varying difficulty levels.

Experimental results

Research questions

  • RQ1Can GPT-4 generate high-quality, curriculum-aligned, and progressively challenging test questions for Grade 9 mathematics?
  • RQ2How does GPT-3.5’s performance in answering complex questions improve after being exposed to questions and feedback from GPT-4?
  • RQ3To what extent can the interaction between GPT-4 (as teacher) and GPT-3.5 (as student) simulate active learning dynamics in AI systems?
  • RQ4How do LLMs adapt their question generation based on feedback and perceived difficulty, reflecting principles of active learning?
  • RQ5What are the comparative strengths and limitations of GPT-3.5 and GPT-4 in generating personalized educational content?

Key findings

  • GPT-4 demonstrated superior ability to generate precise, challenging, and curriculum-aligned test questions across Grade 9 mathematics topics.
  • GPT-3.5 showed notable improvement in handling complex problems after receiving targeted instruction and feedback from GPT-4, indicating adaptive learning potential.
  • The iterative feedback loop between GPT-4 and GPT-3.5 effectively simulated active learning, with question difficulty and content adjusted based on response quality.
  • The system’s performance aligned with uncertainty sampling principles, where GPT-3.5 improved most on harder questions, suggesting engagement with cognitive challenge.
  • GPT-4’s role as a 'teacher' significantly enhanced the quality and depth of questions, outperforming GPT-3.5 in both complexity and accuracy.
  • The results suggest that LLMs can be structured to emulate human teaching and learning dynamics, offering a scalable model for personalized education.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.