Skip to main content
QUICK REVIEW

[论文解读] Designing a realistic peer-like embodied conversational agent for supporting children's storytelling

Zhixin Li, Ying Xu|arXiv (Cornell University)|Apr 19, 2023
AI in Service Interactions被引用 4
一句话总结

本文提出STARie,一种类同龄人的人形对话代理(ECA),利用GPT-3、实时语音克隆、VOCA和FLAME技术,模拟出具有逼真面部动画和声音的儿童风格说书人。该系统旨在通过协作式讲故事提升儿童的叙事能力,同时解决隐私、适龄性、性别表征以及恐怖谷效应等方面的伦理问题。

ABSTRACT

Advances in artificial intelligence have facilitated the use of large language models (LLMs) and AI-generated synthetic media in education, which may inspire HCI researchers to develop technologies, in particular, embodied conversational agents (ECAs) to simulate the kind of scaffolding children might receive from a human partner. In this paper, we will propose a design prototype of a peer-like ECA named STARie that integrates multiple AI models - GPT-3, Speech Synthesis (Real-time Voice Cloning), VOCA (Voice Operated Character Animation), and FLAME (Faces Learned with an Articulated Model and Expressions) that aims to support narrative production in collaborative storytelling, specifically for children aged 4-8. However, designing a child-centered ECA raises concerns about age appropriateness, children privacy, gender choices of ECAs, and the uncanny valley effect. Thus, this paper will also discuss considerations and ethical concerns that must be taken into account when designing such an ECA. This proposal offers insights into the potential use of AI-generated synthetic media in child-centered AI design and how peer-like AI embodiment may support children extquotesingle s storytelling.

研究动机与目标

  • 设计一种类同龄人的人形对话代理(ECA),以支持4至8岁儿童的协作式讲故事。
  • 探究类儿童外貌、声音及实时面部动画如何影响儿童的参与度与叙事发展。
  • 识别并解决在面向儿童的应用中部署AI生成的合成媒体所面临的伦理挑战。
  • 探索类同龄人ECA相较于成人风格ECA在促进叙事能力与情感参与方面的优势。

提出的方法

  • 整合GPT-3以生成自然语言,确保在讲故事过程中产生符合语境且适合儿童的回应。
  • 采用实时语音克隆技术,基于儿童语音数据训练,生成逼真的儿童声音。
  • 使用VOCA(语音驱动角色动画)技术,实现实时同步口语与嘴唇及面部动作。
  • 应用FLAME(带关节结构与表情的面部模型)生成逼真且富有表现力的面部动画。
  • 设计STARie为类儿童女性外貌(年龄约8岁),以模拟同龄人互动,增强亲和力。
  • 集成情感响应机制,如共情式面部表情(例如悲伤、喜悦),以促进儿童进行更深层次的叙事反思。
Figure 1. (a) STARie, the embodied conversational agent, tells a story with a child-like voice, appearance and real-time lip-sync and full-face animation, (b) STARie responds to the child’s previous story with positive feedback and joyful facial expressions, (c) STARie empathizes with the child’s re
Figure 1. (a) STARie, the embodied conversational agent, tells a story with a child-like voice, appearance and real-time lip-sync and full-face animation, (b) STARie responds to the child’s previous story with positive feedback and joyful facial expressions, (c) STARie empathizes with the child’s re

实验结果

研究问题

  • RQ1哪些设计特征对于类同龄人ECA有效支持儿童协作式讲故事与叙事发展至关重要?
  • RQ2在儿童参与度、叙事复杂度与情感反应方面,类儿童ECA与成人风格ECA相比有何差异?
  • RQ3在创建以儿童为中心的ECA时,若其模仿儿童声音与外貌,可能带来哪些关键伦理风险?
  • RQ4如何在面向儿童的AI代理中缓解隐私、适龄性、性别表征以及恐怖谷效应等问题?

主要发现

  • STARie通过整合类儿童声音、外貌与实时面部动画,显著增强了儿童在讲故事过程中对社交存在的感知与参与度。
  • 共情式回应(如悲伤或喜悦的面部表情)可促使儿童反思并扩展其叙事内容,从而支持叙事支架的构建。
  • 儿童可能难以区分物理互动与虚拟互动,凸显了在ECA中精心设计非语言线索的重要性。
  • 在未经过微调的情况下使用大型语言模型(如GPT-3)可能生成不当或有偏见的内容,因此需要内容过滤或针对儿童的微调。
  • 收集与存储儿童语音数据引发重大隐私担忧,尤其是考虑到儿童对数据使用及其长期影响的理解有限。
  • 虽然9岁以下儿童对恐怖谷效应的敏感度可能较低,但若未经过仔细校准,过于逼真却不够完美的动画仍可能引起困惑或降低年轻用户的参与度。
Figure 2. The pipeline that may allow STARie to generate realistic and immediate child-like voices, speech styles, facial expressions
Figure 2. The pipeline that may allow STARie to generate realistic and immediate child-like voices, speech styles, facial expressions

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。