[论文解读] Distributed Microphone Speech Enhancement based on Deep Learning
本文提出三种基于深度神经网络(DNN)的语音增强(SE)系统,用于在扩散噪声环境下的分布式麦克风阵列。DNN–C 系统采用逐通道 DNN 增强后,再通过第二个 DNN 进行融合,其语音质量与可懂度最高,在 TMHINT 数据集上相比单通道(DNN–S)和集中式融合(DNN–F)方法,SSNRI 提升达 16.729 dB,STOI 得分为 0.770。
Speech-related applications deliver inferior performance in complex noise environments. Therefore, this study primarily addresses this problem by introducing speech-enhancement (SE) systems based on deep neural networks (DNNs) applied to a distributed microphone architecture, and then investigates the effectiveness of three different DNN-model structures. The first system constructs a DNN model for each microphone to enhance the recorded noisy speech signal, and the second system combines all the noisy recordings into a large feature structure that is then enhanced through a DNN model. As for the third system, a channel-dependent DNN is first used to enhance the corresponding noisy input, and all the channel-wise enhanced outputs are fed into a DNN fusion model to construct a nearly clean signal. All the three DNN SE systems are operated in the acoustic frequency domain of speech signals in a diffuse-noise field environment. Evaluation experiments were conducted on the Taiwan Mandarin Hearing in Noise Test (TMHINT) database, and the results indicate that all the three DNN-based SE systems provide the original noise-corrupted signals with improved speech quality and intelligibility, whereas the third system delivers the highest signal-to-noise ratio (SNR) improvement and optimal speech intelligibility.
研究动机与目标
- 为在复杂噪声环境中利用分布式麦克风阵列解决语音质量与可懂度下降的问题。
- 研究三种不同 DNN 架构在扩散噪声场中多通道语音增强中的有效性。
- 比较单通道、集中式融合以及两级融合 DNN 结构在分布式麦克风 SE 中的表现。
- 基于 TMHINT 数据库中的客观指标(如 STOI 和 SSNRI)评估性能。
提出的方法
- DNN–S 系统在每个麦克风通道上分别应用一个 DNN,以实现独立的降噪处理。
- DNN–F 系统将所有含噪录音传输至融合中心,在该中心由一个单一 DNN 处理组合后的多通道输入。
- DNN–C 系统首先对每个通道进行本地 DNN 增强,然后通过第二个 DNN 融合增强后的输出,实现最终信号重建。
- 所有系统均在频域中运行,使用在 TMHINT 数据集上、各种噪声条件下的数据进行训练。
- DNN 的训练目标是最小化谱失真,评估基于 STOI(短时客观可懂度)和 SSNRI(信号到子空间噪声比改善)指标。
- 该架构利用空间多样性与深度学习,提升在扩散噪声环境下的语音质量与可懂度。
实验结果
研究问题
- RQ1在语音增强中,每个通道独立进行 DNN 处理的分布式麦克风阵列(DNN–S)与集中式 DNN 处理(DNN–F)相比表现如何?
- RQ2结合本地增强与融合 DNN 阶段的两级 DNN 架构(DNN–C)是否优于单级或集中式方法?
- RQ3麦克风位置与通道特定噪声对扩散噪声环境中基于 DNN 的 SE 性能有何影响?
- RQ4在不同噪声类型与信噪比(SNR)水平下,基于 DNN 的系统在 STOI 和 SSNRI 等客观指标上的表现如何比较?
主要发现
- DNN–C 系统实现了最高的平均 SSNRI 改善值 16.729 dB,显著优于 DNN–S(11.464 dB)和 DNN–F(12.760 dB)。
- DNN–C 系统取得了最佳的 STOI 得分 0.770,超过 DNN–F(0.764)和 DNN–S(0.678),表明其语音可懂度更优。
- DNN–S 系统在各通道间表现不一致,其中 DNN–S1 中的通道 m1 的 STOI(0.670)甚至低于输入信号(0.672),可能由于过拟合所致。
- 时频谱图分析表明,DNN–C 在保留更清晰的谱结构并更有效地抑制噪声方面优于 DNN–F,尤其在高频区域表现更佳。
- DNN–C 系统的时频谱图与干净参考信号最为相似,如图 7(d) 所示的视觉确认。
- 结果证实,将本地增强与融合 DNN 阶段结合,可在分布式麦克风 SE 中实现更优的噪声抑制与可懂度增益。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。