Skip to main content
QUICK REVIEW

[论文解读] Speaker Identification by GMM based i Vector

Soumen Kanrar|arXiv (Cornell University)|Apr 12, 2017
Speech Recognition and Synthesis参考文献 9被引用 3
一句话总结

本文提出了一种基于高斯混合模型(GMM)的i-vector框架用于说话人识别,通过降维和余弦距离评分提升预测准确率。通过使用高斯混合模型建模说话人特征,并应用i-vector自适应以应对信道差异,该方法在多种录音条件和语言环境下实现了更可靠的说话人识别,相较于归一化得分基线在多通道、多语言数据的仿真中表现更优。

ABSTRACT

Speaker Identification process is to identify a particular vocal cord from a set of existing speakers. In the speaker identification processes, unknown speaker voice sample targets each of the existing speakers present in the system and gives a predication. The predication may be more than one existing known speaker voice and is very close to the unknown speaker voice. The model is a Gaussian mixture model built by the extracted acoustic feature vectors from voice. The i-vector based dimension compression mapping function of the channel depended speaker, and super vector give better predicted scores according to cosine distance scoring associated with the order pair of speakers. In the order pair, the first coordinate is the unknown speaker i.e. test speaker, and the second coordinates is the existing known speaker i.e. target speaker. This paper presents the enhancement of the prediction based on i- vector in compare to the normalized set of predicted score. In the simulation, known speaker voices are collected through different channels and in different languages. In the testing, the GMM voice models, and GMM based i-Vector speaker voice models of the known speakers are used among the numbers of clusters in the test data set.

研究动机与目标

  • 在存在信道变化和多样化录音条件的情况下,提升说话人识别的准确率。
  • 探究i-vector自适应在降维的同时保留区分性说话人特征的有效性。
  • 评估余弦距离评分在i-vector超向量上用于说话人识别的性能。
  • 测试基于GMM的i-vector系统在多种语言和录音信道下的鲁棒性。
  • 与归一化得分基线相比,比较所提方法在预测可靠性方面的表现。

提出的方法

  • 系统使用从已知说话人语音样本中提取的声学特征向量训练高斯混合模型(GMMs)。
  • 从GMM中提取i-vectors,将高维说话人表征压缩至低维空间,以减少与信道相关的变异性。
  • 在i-vector映射函数中应用信道补偿技术,以增强对信道效应的鲁棒性。
  • 通过计算测试说话人i-vector(第一维)与已知说话人i-vector(第二维)之间的余弦距离进行说话人验证。
  • 该方法采用由i-vectors构建的超向量,以表示每个说话人的模型,实现高效比较。
  • 使用多信道和多语言测试集对框架进行评估,以检验其泛化能力和鲁棒性。

实验结果

研究问题

  • RQ1基于i-vector的降维在信道变化条件下如何提升说话人识别性能?
  • RQ2与归一化得分基线相比,对i-vector超向量使用余弦距离评分在多大程度上提升了预测准确率?
  • RQ3基于GMM的i-vector系统在不同录音信道和语言下的鲁棒性如何?
  • RQ4在多种声学环境下,该方法能否可靠地从多个已知说话人中识别出未知说话人?
  • RQ5i-vector自适应对说话人模型的紧凑性和区分能力有何影响?

主要发现

  • 与归一化得分基线相比,基于GMM的i-vector方法显著提升了说话人识别的准确率。
  • 该系统在测试数据集中表现出对多种录音信道和多样化语言的鲁棒性能。
  • i-vector降维有效减少了与信道相关的变异性,增强了模型的泛化能力。
  • 在i-vector超向量上使用余弦距离评分,相比其他评分方法,能产生更可靠、更具区分性的预测结果。
  • 该方法在多说话人识别任务中实现了更高的预测置信度和更低的错误率。
  • 仿真结果证实,i-vector自适应增强了系统区分相似说话人声音的能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。