Skip to main content
QUICK REVIEW

[论文解读] Pathological speech detection using x-vector embeddings

Catarina Botelho, Francisco Teixeira|arXiv (Cornell University)|Mar 2, 2020
Voice and Speech Disorders被引用 8
一句话总结

本研究评估了x-vector嵌入作为病理语音检测通用特征提取方法的适用性,结果表明其在使用欧洲葡萄牙语和西班牙葡萄牙语语音语料库检测帕金森病和阻塞性睡眠呼吸暂停时,性能优于基于知识的特征和i-vector,尤其在领域不匹配条件下表现更优。

ABSTRACT

The potential of speech as a non-invasive biomarker to assess a speaker's health has been repeatedly supported by the results of multiple works, for both physical and psychological conditions. Traditional systems for speech-based disease classification have focused on carefully designed knowledge-based features. However, these features may not represent the disease's full symptomatology, and may even overlook its more subtle manifestations. This has prompted researchers to move in the direction of general speaker representations that inherently model symptoms, such as Gaussian Supervectors, i-vectors and, x-vectors. In this work, we focus on the latter, to assess their applicability as a general feature extraction method to the detection of Parkinson's disease (PD) and obstructive sleep apnea (OSA). We test our approach against knowledge-based features and i-vectors, and report results for two European Portuguese corpora, for OSA and PD, as well as for an additional Spanish corpus for PD. Both x-vector and i-vector models were trained with an out-of-domain European Portuguese corpus. Our results show that x-vectors are able to perform better than knowledge-based features in same-language corpora. Moreover, while x-vectors performed similarly to i-vectors in matched conditions, they significantly outperform them when domain-mismatch occurs.

研究动机与目标

  • 评估x-vector嵌入作为病理语音检测通用特征提取方法的适用性。
  • 比较x-vector与基于知识的特征及i-vector在检测帕金森病(PD)和阻塞性睡眠呼吸暂停(OSA)方面的表现。
  • 使用域外训练数据评估模型在领域不匹配情况下的鲁棒性。
  • 在多个多语言语音语料库(包括欧洲葡萄牙语和西班牙葡萄牙语数据集)上验证性能。

提出的方法

  • 在域外的欧洲葡萄牙语语音语料库上训练x-vector模型,以学习通用的说话人表征。
  • 该方法利用深度神经网络从原始语音中提取固定维数的嵌入,捕捉与说话人和疾病相关的特征。
  • 将基于知识的特征(如韵律和谱特征)以及i-vector作为对比基线。
  • 在两个欧洲葡萄牙语语料库上对帕金森病和阻塞性睡眠呼吸暂停进行模型评估,在一个西班牙语语料库上对帕金森病进行评估。
  • 在相同语言(匹配)和不同语言(跨域)两种条件下评估性能。
  • 在提取的嵌入上使用标准机器学习模型进行分类。

实验结果

研究问题

  • RQ1x-vector嵌入能否有效检测帕金森病和阻塞性睡眠呼吸暂停的病理语音?
  • RQ2在同语言检测任务中,x-vector与基于知识的特征相比表现如何?
  • RQ3在领域不匹配条件下,x-vector相对于i-vector的表现如何?
  • RQ4使用域外训练数据是否会影响x-vector在病理语音检测中的泛化能力?

主要发现

  • 在帕金森病和阻塞性睡眠呼吸暂停的同语言检测任务中,x-vector优于基于知识的特征。
  • 在同语言条件下,x-vector与i-vector性能相当。
  • 当训练集与测试集之间存在领域不匹配时,x-vector显著优于i-vector。
  • x-vector学习到的通用说话人表征即使在训练数据来自无关语音领域时,也能有效捕捉与疾病相关的语音模式。
  • 结果表明,x-vector在病理语音检测中比i-vector对领域变化更具鲁棒性。
  • 该方法在语言间表现出强迁移能力,成功将预训练于葡萄牙语的模型应用于西班牙语帕金森病语料库。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。