[논문 리뷰] Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
본 논문은 VR 교육 에이전트를 위한 동적 다중모달 표현(음성 및 제스처)을 생성하는 LLM 기반 방법을 제시하며, 학습자 경험과 참여도가 향상됨을 보인다.
In virtual reality (VR) educational scenarios, Pedagogical agents (PAs) enhance immersive learning through realistic appearances and interactive behaviors. However, most existing PAs rely on static speech and simple gestures. This limitation reduces their ability to dynamically adapt to the semantic context of instructional content. As a result, interactions often lack naturalness and effectiveness in the teaching process. To address this challenge, this study proposes a large language model (LLM)-driven multimodal expression generation method that constructs semantically sensitive prompts to generate coordinated speech and gesture instructions, enabling dynamic alignment between instructional semantics and multimodal expressive behaviors. A VR-based PA prototype was developed and evaluated through user experience-oriented subjective experiments. Results indicate that dynamically generated multimodal expressions significantly enhance learners' perceived learning effectiveness, engagement, and intention to use, while effectively alleviating feelings of fatigue and boredom during the learning process. Furthermore, the combined dynamic expression of speech and gestures notably enhances learners' perceptions of human-likeness and social presence. The findings provide new insights and design guidelines for building more immersive and naturally expressive intelligent PAs.
연구 동기 및 목표
- VR 교육 에이전트에서 적응적이고 의미적으로 정렬된 다중모달 표현의 필요성을 동기를 부여한다.
- 교육적 의미에 따라 조정된 음성 및 제스처 지시를 생성하는 방법을 개발한다.
- 사용자 경험을 평가하기 위한 VR 기반 교육 에이전트 프로토타입을 구축한다.
- 동적 다중모달 표현이 인지된 학습 효과, 몰입, 피로 및 사회적 존재감에 어떤 영향을 미치는지 평가한다.
제안 방법
- 음성 및 제스처 생성을 조정하기 위한 의미적으로 민감한 프롬프트를 제안한다.
- 교육적 의미를 다중모달 표현과 일치시키기 위해 LLM 주도 파이프라인을 사용한다.
- VR 기반 교육 에이전트 프로토타입을 구현한다.
- 영향을 평가하기 위한 사용자 경험 중심의 주관적 실험을 수행한다.
- 인지된 학습 효과, 몰입, 피로 및 사회적 존재감에 대한 영향을 분석한다.
실험 결과
연구 질문
- RQ1정적 표현과 비교하여 동적으로 생성된 다중모달 표현이 인지된 학습 효과를 향상시키는가?
- RQ2조정된 음성 및 제스처 표현이 학습자의 몰입과 에이전트 사용 의사에 영향을 미치는가?
- RQ3VR 학습 중 피로와 권태를 완화하는가?
- RQ4말-제스처 결합 표현이 인지된 인간성 및 사회적 존재감에 어떤 영향을 미치는가?
주요 결과
- 동적으로 생성된 다중모달 표현은 학습자의 인지된 학습 효과, 몰입 및 사용 의도를 크게 향상시킨다.
- 동적 표현은 학습 중 피로와 권태를 완화하는 데 도움이 된다.
- 음성 및 제스처의 결합된 동적 표현은 인간성 및 사회적 존재감에 대한 인식을 향상시킨다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.