Skip to main content
QUICK REVIEW

[论文解读] Stuttering Speech Disfluency Prediction using Explainable Attribution Vectors of Facial Muscle Movements

Arun Das, Jeffrey R. Mock|arXiv (Cornell University)|Oct 2, 2020
Stuttering Research and Treatment参考文献 28被引用 4
一句话总结

本研究提出了一种可解释人工智能(XAI)增强的卷积神经网络(CNN),通过分析发音前的面部肌肉运动,预测成年人口吃者(AWS)即将发生的口吃性不流畅。利用面部表面肌电图(EMG)提取的行动单元(AUs),该模型在不流畅言语前识别出上半部分面部(AU6:提上唇肌)和下半部分面部(AU14:酒窝肌)肌肉活动增强,实现了高精度预测,并通过可解释的归因向量提供模型决策依据。

ABSTRACT

Speech disorders such as stuttering disrupt the normal fluency of speech by involuntary repetitions, prolongations and blocking of sounds and syllables. In addition to these disruptions to speech fluency, most adults who stutter (AWS) also experience numerous observable secondary behaviors before, during, and after a stuttering moment, often involving the facial muscles. Recent studies have explored automatic detection of stuttering using Artificial Intelligence (AI) based algorithm from respiratory rate, audio, etc. during speech utterance. However, most methods require controlled environments and/or invasive wearable sensors, and are unable explain why a decision (fluent vs stuttered) was made. We hypothesize that pre-speech facial activity in AWS, which can be captured non-invasively, contains enough information to accurately classify the upcoming utterance as either fluent or stuttered. Towards this end, this paper proposes a novel explainable AI (XAI) assisted convolutional neural network (CNN) classifier to predict near future stuttering by learning temporal facial muscle movement patterns of AWS and explains the important facial muscles and actions involved. Statistical analyses reveal significantly high prevalence of cheek muscles (p<0.005) and lip muscles (p<0.005) to predict stuttering and shows a behavior conducive of arousal and anticipation to speak. The temporal study of these upper and lower facial muscles may facilitate early detection of stuttering, promote automated assessment of stuttering and have application in behavioral therapies by providing automatic non-invasive feedback in realtime.

研究动机与目标

  • 探究成年人口吃者(AWS)发音前的面部肌肉活动是否包含预测未来言语不流畅的信息。
  • 开发一种非侵入式、可解释人工智能(XAI)框架,利用面部肌肉运动模式对即将发生的言语是否流利进行分类。
  • 识别与口吃可能性相关的特定面部行动单元(AUs)及其时间动态特征。
  • 提供可解释的归因向量,说明哪些面部肌肉和时间窗口对模型预测贡献最大。
  • 探索面部肌肉活动作为实时、非侵入性口吃发作生物标志物的潜力,为未来行为治疗反馈系统提供支持。

提出的方法

  • 利用卷积神经网络(CNN)对AWS在言语准备任务(S1-S2范式)期间采集的时间序列面部肌电图(EMG)信号进行训练。
  • 应用逐层显著性反向传播(LRP)生成可解释的归因向量,突出显示对模型预测贡献最大的面部肌肉运动。
  • 聚焦于面部动作编码系统(FACS)中的行动单元(AUs),特别是AU6(提上唇肌)和AU14(酒窝肌),以量化面部肌肉活动。
  • 处理来自上半部分(脸颊)和下半部分(嘴唇)面部区域的EMG数据,捕捉发声前的运动模式。
  • 在单词对(WG)和非单词对(CW)任务之间比较归因模式,以评估对言语内容的鲁棒性。
  • 采用统计分析(p < 0.005)验证AU6和AU14在预测不流畅性中的显著性。

实验结果

研究问题

  • RQ1发音前的面部肌肉活动能否预测AWS即将发出的言语是流利还是口吃?
  • RQ2哪些特定的面部行动单元(AUs)在流利与不流畅言语试验之间表现出最具区分性的时序模式?
  • RQ3涉及真实词汇(WG)与非词汇(CW)的言语准备任务中,面部肌肉活动模式是否存在差异?这说明了语义内容在其中扮演何种角色?
  • RQ4可解释人工智能(XAI)技术(如LRP)能否可靠地将模型决策归因于特定面部肌肉和时间窗口?
  • RQ5上半部分面部AUs(如AU6)是否反映情绪唤醒,而下半部分面部AUs(如AU14)是否反映即将说话的预期,这是否与时间动态特征相符?

主要发现

  • 在口吃试验中,上半部分面部肌肉AU6(提上唇肌)的激活显著高于流利试验(p < 0.005),在S1启动后800毫秒达到峰值。
  • 下半部分面部肌肉AU14(酒窝肌)在S2启动前呈现持续上升趋势,并在言语即将开始前迅速升高,其在口吃试验中的归因值显著更高(p < 0.005)。
  • AU6与AU14活动的结合提供了强大的预测能力,其独特的时序特征表明二者具有不同功能:AU6反映唤醒,AU14反映言语准备。
  • 归因模式在单词对(WG)和非单词对(CW)任务中保持一致,表明该预测信号不依赖于语义内容。
  • 该模型仅通过非侵入性面部EMG成功区分了流利与不流畅言语,且借助XAI方法实现了高度可解释性。
  • 发音前的面部肌肉活动编码了与口吃可能性相关的内部脑状态,支持将面部运动作为非侵入性生物标志物的使用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。