[Paper Review] Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions
This study investigates how GPT-3.5 and GPT-4 can generate customized, adaptive math test questions for Grade 9 students using an active learning framework. By positioning GPT-4 as a 'teacher' and GPT-3.5 as a 'student' in an iterative feedback loop, the models simulate real-time learning adaptation, with GPT-4 showing superior question quality and GPT-3.5 demonstrating improved performance on complex problems after instruction, highlighting the potential of LLMs to enhance personalized education.
This study investigates how LLMs, specifically GPT-3.5 and GPT-4, can develop tailored questions for Grade 9 math, aligning with active learning principles. By utilizing an iterative method, these models adjust questions based on difficulty and content, responding to feedback from a simulated 'student' model. A novel aspect of the research involved using GPT-4 as a 'teacher' to create complex questions, with GPT-3.5 as the 'student' responding to these challenges. This setup mirrors active learning, promoting deeper engagement. The findings demonstrate GPT-4's superior ability to generate precise, challenging questions and notable improvements in GPT-3.5's ability to handle more complex problems after receiving instruction from GPT-4. These results underscore the potential of LLMs to mimic and enhance active learning scenarios, offering a promising path for AI in customized education. This research contributes to understanding how AI can support personalized learning experiences, highlighting the need for further exploration in various educational contexts
Motivation & Objective
- To evaluate the effectiveness of GPT-3.5 and GPT-4 in generating personalized, curriculum-aligned test questions for Grade 9 mathematics.
- To simulate an active learning environment by using GPT-4 as a 'teacher' and GPT-3.5 as a 'student' to assess iterative learning and adaptation.
- To examine how LLMs respond to increasing question difficulty and feedback, emulating human-like learning progression.
- To assess the potential of large language models to support personalized education through dynamic, adaptive question generation.
- To contribute insights into the integration of LLMs in educational technology for scalable, student-centered learning experiences.
Proposed method
- An iterative question-generation and feedback loop was implemented, where GPT-3.5 responded to questions crafted by GPT-4, with adjustments based on response quality and difficulty.
- GPT-4 was used as a 'teacher' to generate increasingly complex questions in 'Numbers' and 'Financial Mathematics' for Grade 9, targeting specific cognitive levels.
- GPT-3.5 acted as a 'student' responding to these questions, with its answers evaluated for correctness and depth to inform subsequent question refinement.
- The process mimicked active learning by adjusting question difficulty and content based on the model’s performance, simulating formative assessment.
- A simulated 'student' model provided feedback, enabling the system to iteratively improve question quality and alignment with learning objectives.
- Evaluation focused on accuracy, complexity, and adaptability of questions generated across varying difficulty levels.
Experimental results
Research questions
- RQ1Can GPT-4 generate high-quality, curriculum-aligned, and progressively challenging test questions for Grade 9 mathematics?
- RQ2How does GPT-3.5’s performance in answering complex questions improve after being exposed to questions and feedback from GPT-4?
- RQ3To what extent can the interaction between GPT-4 (as teacher) and GPT-3.5 (as student) simulate active learning dynamics in AI systems?
- RQ4How do LLMs adapt their question generation based on feedback and perceived difficulty, reflecting principles of active learning?
- RQ5What are the comparative strengths and limitations of GPT-3.5 and GPT-4 in generating personalized educational content?
Key findings
- GPT-4 demonstrated superior ability to generate precise, challenging, and curriculum-aligned test questions across Grade 9 mathematics topics.
- GPT-3.5 showed notable improvement in handling complex problems after receiving targeted instruction and feedback from GPT-4, indicating adaptive learning potential.
- The iterative feedback loop between GPT-4 and GPT-3.5 effectively simulated active learning, with question difficulty and content adjusted based on response quality.
- The system’s performance aligned with uncertainty sampling principles, where GPT-3.5 improved most on harder questions, suggesting engagement with cognitive challenge.
- GPT-4’s role as a 'teacher' significantly enhanced the quality and depth of questions, outperforming GPT-3.5 in both complexity and accuracy.
- The results suggest that LLMs can be structured to emulate human teaching and learning dynamics, offering a scalable model for personalized education.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.