[Paper Review] Teaching Robots to Span the Space of Functional Expressive Motion
This paper proposes a method for teaching robots to generate functional, expressive motions by learning a shared latent space of Valence-Arousal-Dominance (VAD) from user feedback. Instead of training separate cost functions for each emotion, the robot learns how trajectories map to VAD, enabling efficient generalization to unseen emotions and natural language-based emotion specification, with user studies showing successful emotive style transfer in under 30 minutes of labeling.
Our goal is to enable robots to perform functional tasks in emotive ways, be it in response to their users' emotional states, or expressive of their confidence levels. Prior work has proposed learning independent cost functions from user feedback for each target emotion, so that the robot may optimize it alongside task and environment specific objectives for any situation it encounters. However, this approach is inefficient when modeling multiple emotions and unable to generalize to new ones. In this work, we leverage the fact that emotions are not independent of each other: they are related through a latent space of Valence-Arousal-Dominance (VAD). Our key idea is to learn a model for how trajectories map onto VAD with user labels. Considering the distance between a trajectory's mapping and a target VAD allows this single model to represent cost functions for all emotions. As a result 1) all user feedback can contribute to learning about every emotion; 2) the robot can generate trajectories for any emotion in the space instead of only a few predefined ones; and 3) the robot can respond emotively to user-generated natural language by mapping it to a target VAD. We introduce a method that interactively learns to map trajectories to this latent space and test it in simulation and in a user study. In experiments, we use a simple vacuum robot as well as the Cassie biped.
Motivation & Objective
- Enable robots to perform functional tasks in emotionally expressive ways, adapting motion to user emotions or task confidence.
- Address the inefficiency and lack of generalization in prior methods that train independent cost functions for each predefined emotion.
- Leverage the latent structure of emotions through the Valence-Arousal-Dominance (VAD) space to unify emotion modeling.
- Allow users to teach personalized emotive styles via intuitive labeling, including natural language input mapped to VAD.
- Enable robots to generate expressive motion for any emotion in the VAD space, including those not explicitly trained on.
Proposed method
- Learn a trajectory-to-VAD mapping model using interactive user feedback, where users label trajectories with VAD scores or natural language.
- Use pre-trained language models fine-tuned to predict VAD values from natural language input, enabling emotion inference from phrases.
- Optimize robot motion using a cost function that minimizes the L2 distance between the trajectory’s predicted VAD and a target VAD.
- Iteratively retrain the VAD mapping model using new user labels to improve alignment with human perception of emotion.
- Apply the method in simulation (vacuum robot and Cassie biped) and in a real user study with human participants.
- Support both direct VAD labeling and language-based emotion specification during training and inference.
Experimental results
Research questions
- RQ1Can a single, shared VAD-based model generalize across multiple emotions and reduce the need for per-emotion data collection?
- RQ2Can users efficiently teach personalized emotive styles to robots using intuitive labeling within 30 minutes?
- RQ3Can natural language input be effectively mapped to VAD to enable responsive emotive behavior without explicit emotion labeling?
- RQ4To what extent can the robot generate motion that is perceived as expressing the intended emotion, even when that emotion was not explicitly trained on?
- RQ5How does joint learning of the VAD space improve efficiency and performance compared to training independent cost functions per emotion?
Key findings
- Users successfully taught personalized emotive styles to robots in 30–40 minutes of labeling, demonstrating the method’s practicality.
- In a user study, participants perceived the robot’s intended emotion at a rate significantly higher than random chance (p < 0.05), supporting effective emotion recognition.
- Top-1 accuracy varied across emotions, indicating reasonable but not perfect alignment with target emotions, suggesting robustness over precision.
- Robots trained by different users developed distinct yet justifiable behaviors for the same emotion, showing personalized style learning.
- The method enabled generation of expressive motion for emotions not explicitly queried during training, demonstrating generalization across the VAD space.
- Qualitative results showed that natural language phrases like "Great weather today!" could be mapped to VAD and used to generate appropriate expressive motion.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.