Skip to main content
QUICK REVIEW

[论文解读] Machine learning for the recognition of emotion in the speech of couples in psychotherapy using the Stanford Suppes Brain Lab Psychotherapy Dataset

Colleen Crangle, Rui Wang|arXiv (Cornell University)|Jan 14, 2019
Emotion and Mood Recognition参考文献 21被引用 8
一句话总结

本研究利用机器学习技术,基于斯坦福苏普斯脑实验室心理治疗数据集,对夫妻在自然对话中的情绪(愤怒、悲伤、喜悦、紧张和中性)进行识别。采用滤波器组声学特征与随机森林模型,说话人特定的模型最高达到95%的准确率,表明在存在类别不平衡和非脚本化对话的现实情境下,仍能实现高性能的情绪语音识别。

ABSTRACT

The automatic recognition of emotion in speech can inform our understanding of language, emotion, and the brain. It also has practical application to human-machine interactive systems. This paper examines the recognition of emotion in naturally occurring speech, where there are no constraints on what is said or the emotions expressed. This task is more difficult than that using data collected in scripted, experimentally controlled settings, and fewer results are published. Our data come from couples in psychotherapy. Video and audio recordings were made of three couples (A, B, C) over 18 hour-long therapy sessions. This paper describes the method used to code the audio recordings for the four emotions of Anger, Sadness, Joy and Tension, plus Neutral, also covering our approach to managing the unbalanced samples that a naturally occurring emotional speech dataset produces. Three groups of acoustic features were used in our analysis: filter-bank, frequency, and voice-quality features. The random forests model classified the features. Recognition rates are reported for each individual, the result of the speaker-dependent models that we built. In each case, the best recognition rates were achieved using the filter-bank features alone. For Couple A, these rates were 90% for the female and 87% for the male for the recognition of three emotions plus Neutral. For Couple B, the rates were 84% for the female and 78% for the male for the recognition of all four emotions plus Neutral. For Couple C, a rate of 88% was achieved for the female for the recognition of the four emotions plus Neutral and 95% for the male for three emotions plus Neutral. For pairwise recognition, the rates ranged from 76% to 99% across the three couples. Our results show that couple therapy is a rich context for the study of emotion in naturally occurring speech.

研究动机与目标

  • 开发并评估用于自动识别夫妻心理治疗中自然发生、非脚本化语音情绪的机器学习模型。
  • 应对在临床会话中收集的真实情绪语音数据集中类别不平衡的挑战。
  • 探究不同声学特征集(滤波器组、频谱频率、语音质量)在治疗语音情绪识别中的有效性。
  • 通过分析多个夫妻的说话人特定性能,理解个体在情绪表达与识别中的差异性。
  • 证明利用真实临床数据进行心理健康与人机交互应用中情绪识别的可行性。

提出的方法

  • 从斯坦福苏普斯脑实验室心理治疗数据集中,收集三对夫妻(A、B、C)共18次治疗会话的音频与视频记录。
  • 依据临床标注标准,手动标注语音片段的五类情绪:愤怒、悲伤、喜悦、紧张和中性。
  • 提取三类声学特征:滤波器组能量、频谱频率特征以及语音质量特征(如抖动、闪烁)。
  • 在说话人特定的数据上训练随机森林分类器,以建模个体的情绪表达模式。
  • 应用类别重加权与采样技术,减轻数据集中情绪分布不平衡的影响。
  • 使用标准指标(如准确率)评估模型性能,并对每位说话人及成对比较进行独立分析。

实验结果

研究问题

  • RQ1机器学习模型能否在夫妻心理治疗的自然发生、非脚本化语音中实现高准确率的情绪识别?
  • RQ2在临床语音中,不同声学特征集(滤波器组、频谱频率、语音质量)在情绪识别中的性能表现如何比较?
  • RQ3说话人特定建模在非脚本化治疗对话中,能在多大程度上提升情绪识别的准确率?
  • RQ4真实临床会话中的类别不平衡与情绪可变性,如何影响模型性能与泛化能力?
  • RQ5在真实心理治疗情境下,不同夫妻及个体说话人的情绪识别准确率范围是多少?

主要发现

  • 仅使用滤波器组特征时,准确率最高:在夫妻A中,女性对三种情绪加中性情绪的识别准确率为90%,男性为87%。
  • 在夫妻B中,女性对四种情绪加中性情绪的识别准确率为84%,男性为78%。
  • 在夫妻C中,女性对四种情绪加中性情绪的识别准确率为88%,男性对三种情绪加中性情绪的识别准确率为95%。
  • 跨夫妻的成对情绪识别准确率在76%至99%之间,表明模型在个体间具有较强的泛化潜力。
  • 说话人特定模型显著优于通用模型,凸显了在真实情绪语音识别中个体化建模的重要性。
  • 本研究证明,夫妻心理治疗中非脚本化、临床记录的语音是训练高准确率情绪识别系统可行且丰富的数据来源。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。