[Paper Review] Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
The paper presents an LLM-driven method to generate dynamic multimodal expressions (speech and gestures) for VR pedagogical agents, showing improved learner experience and engagement.
In virtual reality (VR) educational scenarios, Pedagogical agents (PAs) enhance immersive learning through realistic appearances and interactive behaviors. However, most existing PAs rely on static speech and simple gestures. This limitation reduces their ability to dynamically adapt to the semantic context of instructional content. As a result, interactions often lack naturalness and effectiveness in the teaching process. To address this challenge, this study proposes a large language model (LLM)-driven multimodal expression generation method that constructs semantically sensitive prompts to generate coordinated speech and gesture instructions, enabling dynamic alignment between instructional semantics and multimodal expressive behaviors. A VR-based PA prototype was developed and evaluated through user experience-oriented subjective experiments. Results indicate that dynamically generated multimodal expressions significantly enhance learners' perceived learning effectiveness, engagement, and intention to use, while effectively alleviating feelings of fatigue and boredom during the learning process. Furthermore, the combined dynamic expression of speech and gestures notably enhances learners' perceptions of human-likeness and social presence. The findings provide new insights and design guidelines for building more immersive and naturally expressive intelligent PAs.
Motivation & Objective
- Motivate the need for adaptive, semantically aligned multimodal expressions in VR pedagogical agents.
- Develop a method to generate coordinated speech and gesture instructions guided by instructional semantics.
- Build a VR-based pedagogical agent prototype to evaluate user experience.
- Assess how dynamic multimodal expressions influence perceived learning effectiveness, engagement, fatigue, and social presence.
Proposed method
- Propose semantically sensitive prompts to coordinate speech and gesture generation.
- Use an LLM-driven pipeline to align instructional semantics with multimodal expressions.
- Implement a VR-based pedagogical agent prototype.
- Conduct user experience–oriented subjective experiments to evaluate effects.
- Analyze effects on perceived learning effectiveness, engagement, fatigue, and social presence.
Experimental results
Research questions
- RQ1Does dynamically generated multimodal expression improve perceived learning effectiveness compared to static expressions?
- RQ2Do coordinated speech and gesture expressions increase learners' engagement and intention to use the agent?
- RQ3Can dynamic multimodal expressions alleviate fatigue and boredom during VR learning?
- RQ4How does combined speech-gesture expression affect perceived human-likeness and social presence?
Key findings
- Dynamically generated multimodal expressions significantly enhance learners’ perceived learning effectiveness, engagement, and intention to use.
- Dynamic expressions help alleviate feelings of fatigue and boredom during learning.
- Combined dynamic expression of speech and gestures enhances perceptions of human-likeness and social presence.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.