[Paper Review] Designing a realistic peer-like embodied conversational agent for supporting children's storytelling
This paper proposes STARie, a peer-like embodied conversational agent (ECA) that uses GPT-3, real-time voice cloning, VOCA, and FLAME to simulate a child-like storyteller with lifelike facial animation and voice. It aims to enhance children’s narrative skills through collaborative storytelling, while addressing ethical concerns around privacy, age-appropriateness, gender representation, and the uncanny valley effect.
Advances in artificial intelligence have facilitated the use of large language models (LLMs) and AI-generated synthetic media in education, which may inspire HCI researchers to develop technologies, in particular, embodied conversational agents (ECAs) to simulate the kind of scaffolding children might receive from a human partner. In this paper, we will propose a design prototype of a peer-like ECA named STARie that integrates multiple AI models - GPT-3, Speech Synthesis (Real-time Voice Cloning), VOCA (Voice Operated Character Animation), and FLAME (Faces Learned with an Articulated Model and Expressions) that aims to support narrative production in collaborative storytelling, specifically for children aged 4-8. However, designing a child-centered ECA raises concerns about age appropriateness, children privacy, gender choices of ECAs, and the uncanny valley effect. Thus, this paper will also discuss considerations and ethical concerns that must be taken into account when designing such an ECA. This proposal offers insights into the potential use of AI-generated synthetic media in child-centered AI design and how peer-like AI embodiment may support children extquotesingle s storytelling.
Motivation & Objective
- To design a peer-like embodied conversational agent (ECA) that supports collaborative storytelling in children aged 4–8.
- To investigate how child-like appearance, voice, and real-time facial animation influence children’s engagement and narrative development.
- To identify and address ethical challenges in deploying AI-generated synthetic media for child-centered applications.
- To explore the advantages of peer-like ECAs over adult-like ECAs in fostering narrative competence and emotional engagement.
Proposed method
- Integrates GPT-3 for natural language generation to produce contextually appropriate, child-appropriate responses during storytelling.
- Employs real-time voice cloning to generate authentic child-like voices based on training on child speech data.
- Uses VOCA (Voice Operated Character Animation) to synchronize lip and facial movements with spoken dialogue in real time.
- Applies the FLAME (Faces Learned with an Articulated Model and Expressions) model to generate realistic, expressive facial animations.
- Designs STARie with a child-like female appearance (age ~8) to simulate peer interaction and enhance relatability.
- Incorporates emotional response mechanisms such as empathetic facial expressions (e.g., sadness, joy) to prompt deeper narrative reflection.

Experimental results
Research questions
- RQ1What design features are essential for a peer-like ECA to effectively support children’s collaborative storytelling and narrative development?
- RQ2How does a child-like ECA compare to adult-like ECAs in terms of children’s engagement, narrative complexity, and emotional response?
- RQ3What are the key ethical risks associated with creating child-centered ECAs that mimic children’s voices and appearances?
- RQ4How can privacy, age-appropriateness, gender representation, and the uncanny valley effect be mitigated in child-focused AI agents?
Key findings
- STARie’s integration of child-like voice, appearance, and real-time facial animation enhances perceived social presence and engagement in children during storytelling.
- Empathetic responses—such as facial expressions of sadness or joy—can prompt children to reflect and expand their narratives, supporting narrative scaffolding.
- Children may struggle to distinguish between physical and virtual interactions, highlighting the need for careful design of nonverbal cues in ECAs.
- The use of large language models like GPT-3 without fine-tuning risks generating inappropriate or biased content, necessitating content filtering or child-specific fine-tuning.
- Collecting and storing children’s voice data raises significant privacy concerns, especially given children’s limited understanding of data use and long-term implications.
- The uncanny valley effect may be less pronounced in children under 9, but realistic yet imperfect animations can still confuse or disengage young users if not carefully calibrated.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.