[Paper Review] Localization and Tracking of an Acoustic Source using a Diagonal Unloading Beamforming and a Kalman Filter
This paper presents a low-complexity, high-resolution acoustic source localization and tracking system using diagonal unloading (DU) beamforming for direction-of-arrival (DOA) estimation and a Kalman filter (KF) for temporal smoothing. The framework achieves sub-degree RMSE performance on diverse microphone arrays, including linear, pseudo-spherical, and spherical configurations, across static and dynamic scenarios in the LOCATA challenge dataset.
We present the signal processing framework and some results for the IEEE AASP challenge on acoustic source localization and tracking (LOCATA). The system is designed for the direction of arrival (DOA) estimation in single-source scenarios. The proposed framework consists of four main building blocks: pre-processing, voice activity detection (VAD), localization, tracking. The signal pre-processing pipeline includes the short-time Fourier transform (STFT) of the multichannel input captured by the array and the cross power spectral density (CPSD) matrices estimation. The VAD is calculated with a trace-based threshold of the CPSD matrices. The localization is then computed using our recently proposed diagonal unloading (DU) beamforming, which has low-complexity and high resolution. The DOA estimation is finally smoothed with a Kalman filer (KF). Experimental results on the LOCATA development dataset are reported in terms of the root mean square error (RMSE) for a 7-microphone linear array, the 12-microphone pseudo-spherical array integrated in a prototype head for a humanoid robot, and the 32-microphone spherical array.
Motivation & Objective
- Address the challenge of accurate and robust direction-of-arrival (DOA) estimation in single-source acoustic scenarios using microphone arrays.
- Develop a low-complexity, high-resolution beamforming method suitable for real-time implementation in resource-constrained environments.
- Integrate a Kalman filter to smooth DOA estimates over time, improving tracking stability in noisy or dynamic conditions.
- Evaluate the system across diverse microphone array configurations—linear, pseudo-spherical, and spherical—on the LOCATA development dataset.
- Optimize the signal processing pipeline for practical deployment in applications such as human-robot interaction, teleconferencing, and audio surveillance.
Proposed method
- Apply short-time Fourier transform (STFT) to multichannel audio signals, using a Hann window, 2048-point FFT, and 512-point hop size.
- Estimate cross-power spectral density (CPSD) matrices by averaging over 25 frames to improve spectral coherence and noise robustness.
- Implement a trace-based voice activity detection (VAD) to identify active speech segments, with empirically tuned thresholds per array type.
- Perform DOA estimation using diagonal unloading (DU) beamforming, which enhances resolution by subtracting a trace-scaled identity matrix from the CPSD matrix.
- Compute broadband steered response power (SRP) via incoherent frequency fusion of narrowband DU responses across the 80–8000 Hz band.
- Apply an extended Kalman filter (EKF) to track DOA over time, using state prediction and correction steps with process and measurement noise parameters.
Experimental results
Research questions
- RQ1How does diagonal unloading beamforming compare to conventional beamformers in terms of DOA estimation resolution and computational complexity?
- RQ2To what extent does Kalman filtering improve the temporal stability and accuracy of DOA estimates in dynamic speaker scenarios?
- RQ3How does the system perform across different microphone array geometries—linear, pseudo-spherical, and spherical—under varying acoustic conditions?
- RQ4What is the impact of VAD threshold selection on localization accuracy and robustness to background noise?
- RQ5Can the proposed framework achieve high-accuracy DOA estimation in real-world, non-anechoic environments with moving sources and arrays?
Key findings
- The system achieved an azimuth RMSE of 0.972° on the 7-microphone linear array for static speaker and static array conditions (task 1, recording 3).
- For the moving speaker and static array scenario (task 3), the robot head array achieved an azimuth RMSE of 2.880° and elevation RMSE of 2.807° on recording 3.
- On the most challenging task (task 5, moving speaker and moving array), the eigenmike spherical array achieved an azimuth RMSE of 4.433° and elevation RMSE of 3.100° on recording 1.
- The system demonstrated robustness to array motion and speaker dynamics, with the lowest RMSE values observed in static or low-motion scenarios.
- The VAD threshold was empirically optimized per array type, with values of 200 (linear), 50 (robot head), and 10 (eigenmike), significantly improving detection reliability.
- The Kalman filter effectively reduced DOA estimation jitter, particularly in high-motion and noisy conditions, as evidenced by smoother trajectories in Figures 3 and 4.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.