Skip to main content
QUICK REVIEW

[Paper Review] Speaker Identification by GMM based i Vector

Soumen Kanrar|arXiv (Cornell University)|Apr 12, 2017
Speech Recognition and Synthesis9 references3 citations
TL;DR

This paper proposes a GMM-based i-vector framework for speaker identification that improves prediction accuracy through dimensionality reduction and cosine distance scoring. By modeling speaker characteristics via Gaussian Mixture Models and applying i-vector adaptation to handle channel variations, the method achieves more reliable speaker recognition across diverse recording conditions and languages, outperforming normalized score baselines in simulations with multi-channel, multilingual data.

ABSTRACT

Speaker Identification process is to identify a particular vocal cord from a set of existing speakers. In the speaker identification processes, unknown speaker voice sample targets each of the existing speakers present in the system and gives a predication. The predication may be more than one existing known speaker voice and is very close to the unknown speaker voice. The model is a Gaussian mixture model built by the extracted acoustic feature vectors from voice. The i-vector based dimension compression mapping function of the channel depended speaker, and super vector give better predicted scores according to cosine distance scoring associated with the order pair of speakers. In the order pair, the first coordinate is the unknown speaker i.e. test speaker, and the second coordinates is the existing known speaker i.e. target speaker. This paper presents the enhancement of the prediction based on i- vector in compare to the normalized set of predicted score. In the simulation, known speaker voices are collected through different channels and in different languages. In the testing, the GMM voice models, and GMM based i-Vector speaker voice models of the known speakers are used among the numbers of clusters in the test data set.

Motivation & Objective

  • To improve speaker identification accuracy in the presence of channel variations and diverse recording conditions.
  • To investigate the effectiveness of i-vector adaptation in reducing dimensionality while preserving discriminative speaker features.
  • To evaluate performance of cosine distance scoring on i-vector super-vectors for speaker recognition.
  • To test the robustness of the GMM-based i-vector system across multiple languages and recording channels.
  • To compare the proposed method with normalized score baselines in terms of prediction reliability.

Proposed method

  • The system uses Gaussian Mixture Models (GMMs) trained on acoustic feature vectors extracted from known speaker voice samples.
  • i-vectors are derived from the GMMs to compress high-dimensional speaker representations into lower-dimensional space, reducing channel-dependent variability.
  • A channel compensation technique is applied to the i-vector mapping function to enhance robustness to channel effects.
  • Speaker verification is performed using cosine distance scoring between the test speaker's i-vector (first coordinate) and known speaker i-vectors (second coordinate).
  • The method employs super-vectors constructed from i-vectors to represent each speaker's model for efficient comparison.
  • The framework is evaluated using multi-channel and multilingual test sets to assess generalization and robustness.

Experimental results

Research questions

  • RQ1How does i-vector-based dimensionality reduction improve speaker identification performance under channel variations?
  • RQ2To what extent does cosine distance scoring on i-vector super-vectors enhance prediction accuracy compared to normalized score baselines?
  • RQ3How robust is the GMM-based i-vector system across different recording channels and languages?
  • RQ4Can the proposed method reliably identify unknown speakers among multiple known speakers in diverse acoustic environments?
  • RQ5What is the impact of i-vector adaptation on speaker model compactness and discriminative power?

Key findings

  • The GMM-based i-vector approach significantly improves speaker identification accuracy compared to normalized score baselines.
  • The system demonstrates robust performance across multiple recording channels and diverse languages in the test dataset.
  • i-vector dimensionality compression effectively reduces channel-dependent variability, enhancing model generalization.
  • Cosine distance scoring on i-vector super-vectors yields more reliable and discriminative predictions than alternative scoring methods.
  • The method achieves higher prediction confidence and lower error rates in multi-speaker identification tasks.
  • The simulation results confirm that i-vector adaptation enhances the system's ability to distinguish between similar speaker voices.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.