[Paper Review] Stuttering Speech Disfluency Prediction using Explainable Attribution Vectors of Facial Muscle Movements
This study proposes an explainable AI (XAI)-enhanced convolutional neural network (CNN) that predicts upcoming stuttering disfluencies in adults who stutter (AWS) by analyzing pre-speech facial muscle movements. Using Action Units (AUs) from facial electromyography (EMG), the model identifies heightened activity in upper face (AU6: cheek raiser) and lower face (AU14: dimpler) muscles before disfluent speech, achieving high predictive accuracy with interpretable attribution vectors.
Speech disorders such as stuttering disrupt the normal fluency of speech by involuntary repetitions, prolongations and blocking of sounds and syllables. In addition to these disruptions to speech fluency, most adults who stutter (AWS) also experience numerous observable secondary behaviors before, during, and after a stuttering moment, often involving the facial muscles. Recent studies have explored automatic detection of stuttering using Artificial Intelligence (AI) based algorithm from respiratory rate, audio, etc. during speech utterance. However, most methods require controlled environments and/or invasive wearable sensors, and are unable explain why a decision (fluent vs stuttered) was made. We hypothesize that pre-speech facial activity in AWS, which can be captured non-invasively, contains enough information to accurately classify the upcoming utterance as either fluent or stuttered. Towards this end, this paper proposes a novel explainable AI (XAI) assisted convolutional neural network (CNN) classifier to predict near future stuttering by learning temporal facial muscle movement patterns of AWS and explains the important facial muscles and actions involved. Statistical analyses reveal significantly high prevalence of cheek muscles (p<0.005) and lip muscles (p<0.005) to predict stuttering and shows a behavior conducive of arousal and anticipation to speak. The temporal study of these upper and lower facial muscles may facilitate early detection of stuttering, promote automated assessment of stuttering and have application in behavioral therapies by providing automatic non-invasive feedback in realtime.
Motivation & Objective
- To investigate whether pre-speech facial muscle activity in adults who stutter (AWS) contains predictive information for upcoming speech disfluency.
- To develop a non-invasive, explainable AI (XAI) framework that classifies upcoming speech as fluent or stuttered using facial muscle movement patterns.
- To identify specific facial Action Units (AUs) and their temporal dynamics that correlate with stuttering likelihood.
- To provide interpretable attribution vectors explaining which facial muscles and time windows contribute most to the model's predictions.
- To explore the potential of facial muscle activity as a real-time, non-invasive biomarker for stuttering onset, enabling future behavioral therapy feedback systems.
Proposed method
- Utilizes a convolutional neural network (CNN) trained on time-series facial electromyography (EMG) signals from AWS during a speech preparation task (S1-S2 paradigm).
- Applies Layer-wise Relevance Backpropagation (LRP) to generate explainable attribution vectors that highlight which facial muscle movements contribute most to the model’s prediction.
- Focuses on Action Units (AUs) from the Facial Action Coding System (FACS), specifically AU6 (cheek raiser) and AU14 (dimpler), to quantify facial muscle activity.
- Processes EMG data from upper (cheek) and lower (lip) facial regions to capture pre-speech motor patterns before vocalization onset.
- Compares attribution patterns across two speech tasks: word pair (WG) and non-word pair (CW) conditions to assess robustness to speech content.
- Employs statistical analysis (p < 0.005) to validate the significance of AU6 and AU14 in predicting disfluency.
Experimental results
Research questions
- RQ1Can pre-speech facial muscle activity predict whether an upcoming utterance in AWS will be fluent or stuttered?
- RQ2Which specific facial Action Units (AUs) show the most discriminative temporal patterns between fluent and disfluent speech trials?
- RQ3Do facial muscle activity patterns differ between speech preparation tasks involving real words (WG) versus non-words (CW), and what does this imply about the role of semantic content?
- RQ4Can explainable AI (XAI) techniques like LRP reliably attribute model decisions to specific facial muscles and time windows?
- RQ5Do upper facial AUs (e.g., AU6) reflect emotional arousal, while lower facial AUs (e.g., AU14) reflect anticipation to speak, as suggested by temporal dynamics?
Key findings
- Significantly higher activation in upper facial muscle AU6 (cheek raiser) was observed in stuttered trials compared to fluent ones (p < 0.005), peaking 800 ms after S1 onset.
- Lower facial muscle AU14 (dimpler) showed a sustained increase before S2 onset and a rapid rise just prior to speech, with significantly higher attribution in stuttered trials (p < 0.005).
- The combination of AU6 and AU14 activity provided strong predictive power, with distinct temporal profiles suggesting separate roles: AU6 for arousal and AU14 for anticipation.
- Attribution patterns were consistent across both word pair (WG) and non-word pair (CW) tasks, indicating that the predictive signal is not dependent on semantic content.
- The model successfully distinguished fluent from disfluent speech using only non-invasive facial EMG, with high interpretability via XAI methods.
- Facial muscle activity before speech onset encodes internal brain states related to stuttering likelihood, supporting the use of facial movements as a non-invasive biomarker.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.