[Paper Review] GuideAI: A Real-time Personalized Learning Solution with Adaptive Interventions
GuideAI presents a biosensor-augmented, real-time adaptive learning system that uses multi-modal data to infer cognitive-affective states and deliver personalized interventions across text, image, audio, and video modalities. A preliminary study showed improvements in retention and reduced cognitive load with LLM-based tutoring.
Large Language Models (LLMs) have emerged as powerful learning tools, but they lack awareness of learners' cognitive and physiological states, limiting their adaptability to the user's learning style. Contemporary learning techniques primarily focus on structured learning paths, knowledge tracing, and generic adaptive testing but fail to address real-time learning challenges driven by cognitive load, attention fluctuations, and engagement levels. Building on findings from a formative user study (N=66), we introduce GuideAI, a multi-modal framework that enhances LLM-driven learning by integrating real-time biosensory feedback including eye gaze tracking, heart rate variability, posture detection, and digital note-taking behavior. GuideAI dynamically adapts learning content and pacing through cognitive optimizations (adjusting complexity based on learning progress markers), physiological interventions (breathing guidance and posture correction), and attention-aware strategies (redirecting focus using gaze analysis). Additionally, GuideAI supports diverse learning modalities, including text-based, image-based, audio-based, and video-based instruction, across varied knowledge domains. A preliminary study (N = 25) assessed GuideAI's impact on knowledge retention and cognitive load through standardized assessments. The results show statistically significant improvements in both problem-solving capability and recall-based knowledge assessments. Participants also experienced notable reductions in key NASA-TLX measures including mental demand, frustration levels, and effort, while simultaneously reporting enhanced perceived performance. These findings demonstrate GuideAI's potential to bridge the gap between current LLM-based learning systems and individualized learner needs, paving the way for adaptive, cognition-aware education at scale.
Motivation & Objective
- Identify limitations of current LLM-based tutors in real-time learner-state awareness and multimodal adaptation.
- Propose a biosensor-augmented, closed-loop learning framework that infers cognitive-affective states from multi-modal signals.
- Develop a multi-modal GuideAI system supporting text, image, audio, and video modalities.
- Evaluate GuideAI through formative studies and a preliminary user study to assess learning outcomes and cognitive load.
- Open-source the GuideAI codebase to enable reproducibility and further research.
Proposed method
- Formative study with N=66 to identify learner needs and design implications.
- Three-module GuideAI architecture: Sensor Module (biometric and behavioral data), Processing Module (signal processing), Inference Module (state estimation and interventions).
- Biosensor data streams include eye tracking, HRV, posture, and note-taking behavior; signals are time-synced via Lab Streaming Layer (LSL).
- State inference computes six cognitive dimensions (cognitive load, attention, engagement, understanding, stress, fatigue) from normalized, baseline-adjusted features.
- Interventions are delivered via LLM/VLM/Audio LLM with tone adaptations and modality-specific strategies.
- Intervention thresholds are baseline-relative Z-scores with thresholds |z|≥1.0 for moderate deviation and |z|≥1.5 for pronounced deviation, enforced over a 10s persistence window.

Experimental results
Research questions
- RQ1Can a biosensor-augmented LLM system infer real-time cognitive-affective learner states across multiple modalities?
- RQ2Do real-time, adaptive interventions based on inferred states improve learning outcomes and reduce cognitive load compared to non-personalized baselines?
- RQ3How can interventions be effectively tailored across text, image, audio, and video modalities to maintain learning flow?
- RQ4What is the feasibility of deploying a cross-device, real-time educational platform that aggregates gaze, HRV, posture, and notes data?
- RQ5Is the GuideAI approach scalable and generalizable across varied knowledge domains and learner populations?
Key findings
- Preliminary study (N=25) showed statistically significant improvements in problem solving and recall-based knowledge assessments compared to non-personalized baselines.
- Participants reported reductions in NASA-TLX measures such as mental demand, frustration, and effort, with enhanced perceived performance.
- Biometric and behavioral cues enabled effective, real-time adaptations such as content pacing, complexity modulation, and physiological interventions (e.g., box breathing, posture cues).
- GuideAI supports four learning modalities (text, image, audio, video) with modality-specific interventions that adjust to cognitive load and engagement levels.
- Formative study informed design implications for multi-modal support, real-time adaptation, physiological awareness, and personalized learning paths.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.