[Paper Review] Look at Me When I Talk to You: A Video Dataset to Enable Voice Assistants to Recognize Errors
This paper introduces an open-source video dataset of 21 participants' facial reactions during interactions with a voice assistant, demonstrating that facial expressions alone can signal whether the assistant made an error. By using crowdsourced workers to classify reactions from silent video clips, the study shows promising trends in recognizing voice assistant errors through visual cues, paving the way for self-repair mechanisms in conversational AI.
People interacting with voice assistants are often frustrated by voice assistants' frequent errors and inability to respond to backchannel cues. We introduce an open-source video dataset of 21 participants' interactions with a voice assistant, and explore the possibility of using this dataset to enable automatic error recognition to inform self-repair. The dataset includes clipped and labeled videos of participants' faces during free-form interactions with the voice assistant from the smart speaker's perspective. To validate our dataset, we emulated a machine learning classifier by asking crowdsourced workers to recognize voice assistant errors from watching soundless video clips of participants' reactions. We found trends suggesting it is possible to determine the voice assistant's performance from a participant's facial reaction alone. This work posits elicited datasets of interactive responses as a key step towards improving error recognition for repair for voice assistants in a wide variety of applications.
Motivation & Objective
- To address the lack of error recognition in voice assistants, which often fail to detect their own mistakes during interactions.
- To explore whether facial reactions can serve as reliable indicators of voice assistant errors, enabling automatic detection and self-repair.
- To create and release a publicly available, ethically collected video dataset of user reactions to correct, incorrect, or non-responses from a voice assistant.
- To validate the feasibility of training machine learning models on visual cues alone to recognize when a voice assistant has erred.
- To lay the foundation for future research on error detection, repair, and conversational grounding in voice-enabled systems using real human interaction data.
Proposed method
- Recorded 21 participants interacting with Amazon Alexa in a naturalistic setting, capturing audio and video from the smart speaker’s perspective.
- Extracted and clipped silent video segments showing participants’ facial expressions during responses to Alexa’s answers.
- Classified each video clip into three categories: correct response (YES), incorrect response (NO), or no response (OMIT).
- Employed crowdsourced workers on Amazon Mechanical Turk to label the silent video clips as indicating an error (NO) or no error (YES/OMIT).
- Used the labeled dataset to emulate a machine learning classifier for error recognition based solely on facial expressions.
- Implemented ethical safeguards, including IRB approval, participant consent, and restricted data distribution to researchers with IRB clearance.
Experimental results
Research questions
- RQ1Can facial expressions alone reliably indicate whether a voice assistant has made an error during a conversation?
- RQ2How effective are crowdsourced workers at detecting voice assistant errors from silent video clips of user reactions?
- RQ3What are the most salient facial cues associated with user recognition of incorrect or missing responses from voice assistants?
- RQ4To what extent can visual reaction data be used to train machine learning models for automatic error detection in voice assistants?
- RQ5How can interactive response datasets improve the development of self-repair mechanisms in conversational AI systems?
Key findings
- Crowdsourced workers were able to classify the presence of voice assistant errors with notable consistency when viewing silent video clips of participants’ facial reactions.
- The study found clear trends in facial expressions—such as head shakes, eye rolls, or lack of response—associated with incorrect or missing responses from the assistant.
- Participants exhibited distinct visual reactions to incorrect responses (e.g., head shakes) and non-responses (e.g., silence or gaze shifts), suggesting detectable patterns.
- The dataset demonstrates that facial reactions contain sufficient information to infer the assistant’s performance, even without audio.
- The research validates the feasibility of using visual-only cues for training error-detection models in voice assistants, supporting future real-time, on-device inference.
- Ethical considerations were addressed through IRB approval, informed consent, and restricted data access, ensuring responsible use of sensitive biometric data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.