Skip to main content
QUICK REVIEW

[Paper Review] Auxiliary Function-Based Algorithm for Blind Extraction of a Moving Speaker

Jakub Janský, Zbyněk Koldovský|arXiv (Cornell University)|Feb 28, 2020
Blind Source Separation Techniques4 citations
TL;DR

This paper proposes Block AuxIVE, a novel blind source extraction algorithm based on the Constant Separating Vector (CSV) model for tracking moving speakers in reverberant and noisy environments. By leveraging an auxiliary function-based optimization, the method achieves faster convergence and robust extraction without online tracking, outperforming static models in real-world and CHiME-4 benchmark conditions.

ABSTRACT

Recently, Constant Separating Vector (CSV) mixing model has been proposed for the Blind Source Extraction (BSE) of moving sources. In this paper, we experimentally verify the applicability of CSV in the blind extraction of a moving speaker and propose a new BSE method derived by modifying the auxiliary function-based algorithm for Independent Vector Analysis. Also, a piloted variant is proposed for the method with partially controllable global convergence. The methods are verified under reverberant and noisy conditions using {\color{red} simulated as well as real-world acoustic conditions}. They are also verified within the CHiME-4 speech separation and recognition challenge. The experiments corroborate the applicability of CSV as well as the improved convergence of the proposed algorithms.

Motivation & Objective

  • Address the challenge of blind extraction of a moving speaker in multi-microphone acoustic environments where source positions change over time.
  • Overcome the discontinuity problem in on-line BSE caused by switching between different static mixing models for short intervals.
  • Develop a BSE method that maintains consistent extraction of the target source across source movements without explicit source tracking.
  • Improve convergence speed and global convergence stability compared to gradient-based BSE methods, especially in overdetermined microphone arrays.
  • Verify the method’s robustness in highly reverberant and noisy conditions using simulated and real-world data, including the CHiME-4 challenge dataset.

Proposed method

  • Adopt the Constant Separating Vector (CSV) mixing model, which assumes a single time-invariant separating filter can extract the target source across its movement path.
  • Propose Block AuxIVE, a block-wise BSE algorithm based on auxiliary function optimization, enabling faster convergence than gradient-based counterparts.
  • Utilize the STFT domain to model convoluted mixtures as frequency-wise instantaneous mixtures, enabling efficient joint processing across frequency bins.
  • Apply the auxiliary function method to derive a fixed-point update rule that ensures monotonic convergence of the objective function.
  • Introduce a semi-supervised variant using pilot signals—partially pre-extracted target sources with residual noise—to stabilize global convergence.
  • Use Relative Transfer Function (RTF) estimation for initialization and apply Blind Analytic Normalization (BAN) as a postfilter in comparison systems.

Experimental results

Research questions

  • RQ1Can the CSV model effectively extract a moving speaker without explicit source tracking, and how does it compare to static mixing models in dynamic environments?
  • RQ2Does the proposed auxiliary function-based algorithm (Block AuxIVE) achieve faster convergence and better global convergence than gradient-based BSE methods?
  • RQ3How robust is the Block AuxIVE algorithm to reverberation, noise, and imperfect initialization in real-world and simulated acoustic conditions?
  • RQ4Can the semi-supervised variant with pilot signals ensure stable global convergence even when the pilot contains significant noise and interference?
  • RQ5What is the performance of Block AuxIVE in real-world benchmarks such as the CHiME-4 challenge, particularly for moving speaker scenarios?

Key findings

  • Block AuxIVE converges in only 7 iterations on average, reducing total processing time from 10h30min (BOGIVE w) to 1h2min for the full CHiME-4 dataset.
  • The method achieves a Word Error Rate (WER) of 6.44% on the real test set and 6.46% on the simulated test set, outperforming OverIVA (10.43% and 6.82%) and GEV (8.10% and 5.99%) in some conditions.
  • Block AuxIVE successfully recovers speech during speaker movement, while OverIVA fails due to fixed beamforming, as shown in Fig. 5 where voice vanishes when the speaker moves.
  • The semi-supervised variant with pilot signals achieves stable global convergence even when the pilot contains significant residual noise and interference.
  • The CSV-based approach enables consistent extraction across source movements, avoiding the discontinuity problem inherent in frame-wise static model application.
  • The proposed method achieves performance comparable to trained GEV beamformer without requiring any training data, demonstrating strong generalization in blind settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.