[论文解读] Real-time binaural speech separation with preserved spatial cues
本文提出了一种实时、端到端的多输入多输出(MIMO)时域语音分离网络(MIMO TasNet),通过使用并行编码器和掩码-求和机制,保留了双耳线索——双耳时间差(ITD)和双耳强度差(ILD)。该方法在各种声学条件下(包括噪声和混响)实现了高保真度、低延迟的双耳语音分离,显著提升了语音质量和空间线索的保留效果。
Deep learning speech separation algorithms have achieved great success in improving the quality and intelligibility of separated speech from mixed audio. Most previous methods focused on generating a single-channel output for each of the target speakers, hence discarding the spatial cues needed for the localization of sound sources in space. However, preserving the spatial information is important in many applications that aim to accurately render the acoustic scene such as in hearing aids and augmented reality (AR). Here, we propose a speech separation algorithm that preserves the interaural cues of separated sound sources and can be implemented with low latency and high fidelity, therefore enabling a real-time modification of the acoustic scene. Based on the time-domain audio separation network (TasNet), a single-channel time-domain speech separation system that can be implemented in real-time, we propose a multi-input-multi-output (MIMO) end-to-end extension of TasNet that takes binaural mixed audio as input and simultaneously separates target speakers in both channels. Experimental results show that the proposed end-to-end MIMO system is able to significantly improve the separation performance and keep the perceived location of the modified sources intact in various acoustic scenes.
研究动机与目标
- 解决单通道语音分离方法因丢弃空间线索而导致的局限性,而这些空间线索对声音定位至关重要。
- 开发一种实时、端到端的双耳分离系统,以保留双耳时间差(ITD)和双耳强度差(ILD)。
- 实现与听者无关的低延迟声学场景修改,适用于助听器和增强现实(AR)等应用。
- 评估MIMO TasNet在具有挑战性的环境中的性能,包括噪声和混响条件。
提出的方法
- 将时域音频分离网络(TasNet)扩展为多输入多输出(MIMO)框架,以直接处理双耳混合信号。
- 采用并行编码器,同时从左右麦克风信号中提取跨通道特征。
- 将双通道编码器输出拼接,并通过共享的掩码估计网络生成空间和频谱掩码。
- 应用掩码-求和机制,对每条通道独立执行联合频谱与空间滤波。
- 使用线性解码器从掩码后的编码器输出重建双耳分离语音信号。
- 在部分变体中,将双耳特征(IPD、ILD)作为输入嵌入,以提升线索保留效果。
实验结果
研究问题
- RQ1基于深度学习的MIMO语音分离系统是否能在实现实时性能的同时,保留双耳线索(ITD与ILD)?
- RQ2MIMO TasNet在语音质量与空间线索保留方面,相较于单通道TasNet及其他MIMO变体表现如何?
- RQ3该方法在不同说话者角度和环境条件下,对空间定位准确性的保持程度如何?
- RQ4在添加噪声和房间混响等不利条件下,系统性能表现如何?
主要发现
- 采用并行编码器和掩码-求和机制的MIMO TasNet在混响环境中实现了最高的分离质量(信噪比提升9.4 dB)和最低的ITD误差(5.9 μs)。
- 系统最小延迟低于5 ms,支持可穿戴设备中的实时部署。
- 与单通道TasNet基线相比,MIMO TasNet在所有说话者角度下将ITD误差降低了70%,ILD误差降低了60%。
- 在噪声条件下,采用并行编码器的MIMO TasNet优于所有其他变体,在所有信噪比水平下均保持了更优的语音质量和空间线索保留。
- 信噪比提升与ITD/ILD误差降低之间存在较强相关性(分别为r = -0.77 和 r = -0.85),证实了更优的分离效果可提升空间线索保真度。
- 即使在三说话者混合场景中,MIMO TasNet仍保持显著性能,但因低功率或空间相似的源被误分类,ITD与ILD的保留效果有所下降。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。