[论文解读] Symphony: Localizing Multiple Acoustic Sources with a Single Microphone Array
Symphony 提出了一种使用单个麦克风阵列实现多个声源定位的首种方法,通过基于几何的滤波分离不同路径的信号,并利用相干性聚类将来自同一声源的信号分组。其定位误差中位数为 0.694 米,较最先进方法降低 68%,实现了在消费级设备上无需分布式阵列即可进行高精度、实时的声源定位。
Sound recognition is an important and popular function of smart devices. The location of sound is basic information associated with the acoustic source. Apart from sound recognition, whether the acoustic sources can be localized largely affects the capability and quality of the smart device's interactive functions. In this work, we study the problem of concurrently localizing multiple acoustic sources with a smart device (e.g., a smart speaker like Amazon Alexa). The existing approaches either can only localize a single source, or require deploying a distributed network of microphone arrays to function. Our proposal called Symphony is the first approach to tackle the above problem with a single microphone array. The insight behind Symphony is that the geometric layout of microphones on the array determines the unique relationship among signals from the same source along the same arriving path, while the source's location determines the DoAs (direction-of-arrival) of signals along different arriving paths. Symphony therefore includes a geometry-based filtering module to distinguish signals from different sources along different paths and a coherence-based module to identify signals from the same source. We implement Symphony with different types of commercial off-the-shelf microphone arrays and evaluate its performance under different settings. The results show that Symphony has a median localization error of 0.694m, which is 68% less than that of the state-of-the-art approach.
研究动机与目标
- 通过单个紧凑型麦克风阵列实现多个声源的高精度定位,克服现有系统存在的局限性。
- 解决在密集声学环境中多个声源同时发射时的声源分离与定位难题。
- 消除对分布式麦克风阵列的需求,使该方案适用于智能音箱等消费级智能设备。
- 开发一种对不同麦克风阵列配置和真实世界条件均具有鲁棒性的方法。
- 在最小硬件需求下实现高精度定位,适用于商用现成设备。
提出的方法
- Symphony 采用基于几何的滤波模块,利用麦克风固定的物理空间布局,区分来自不同声源、沿不同路径传播的信号。
- 通过测量麦克风之间独特的时延与振幅关系,建模信号传播特性,并依据其几何特征分离声源。
- 基于相干性的模块通过测量阵列各元件间信号的相似性,识别并聚类来自同一声源的信号。
- 通过结合滤波后的信号与相干性评分,联合估计每个声源的波达方向(DoA)。
- 采用端到端处理流水线,首先分离路径相关的信号分量,再按声源相干性进行分组,实现实时处理。
- 该方法在多种商用麦克风阵列上实现,展现出对不同设备形态的硬件无关性能。
实验结果
研究问题
- RQ1是否仅通过单个麦克风阵列即可实现多个声源的高精度定位,而无需依赖分布式网络?
- RQ2当多个声源的信号沿重叠路径到达单个阵列时,如何实现有效分离?
- RQ3麦克风的几何布局在区分声源特异性信号分量方面起到何种作用?
- RQ4在多声源场景下,基于相干性的聚类如何提升声源定位精度?
- RQ5所提出的方法是否能在真实世界硬件上实现优于最先进方法的定位性能?
主要发现
- Symphony 在各种测试条件下实现了 0.694 米的中位数定位误差,显著优于最先进方法。
- 与先前最先进方法相比,中位数误差降低 68%,表明定位精度获得显著提升。
- 该方法在不同类型商用现成麦克风阵列上均保持一致的性能表现,表明对硬件差异具有鲁棒性。
- 基于几何的滤波能有效分离来自相似方向的多个声源信号。
- 基于相干性的聚类可成功将同一声源的信号分组,从而实现每个声源的精确波达方向估计。
- 系统可实时运行,并可部署于标准智能设备(如 Amazon Alexa),具备实际集成可行性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。