Skip to main content
QUICK REVIEW

[论文解读] Cracking the cocktail party problem by multi-beam deep attractor network

Zhuo Chen, Jinyu Li|arXiv (Cornell University)|Mar 29, 2018
Speech and Audio Processing参考文献 33被引用 3
一句话总结

本文提出了一种多波束深度吸引子网络(MBDAN),通过首先对多通道输入应用固定差分波束成形器,然后在每个波束上独立使用单通道锚定深度吸引子网络(DAN),最后进行最优波束选择,从而提升多说话人语音分离性能。该方法在4人、3人和2人混合语音中分别实现了11.5 dB、11.76 dB和11.02 dB的SIR提升,性能与使用理想MVDR波束成形器的系统相当。

ABSTRACT

While recent progresses in neural network approaches to single-channel speech separation, or more generally the cocktail party problem, achieved significant improvement, their performance for complex mixtures is still not satisfactory. In this work, we propose a novel multi-channel framework for multi-talker separation. In the proposed model, an input multi-channel mixture signal is firstly converted to a set of beamformed signals using fixed beam patterns. For this beamforming, we propose to use differential beamformers as they are more suitable for speech separation. Then each beamformed signal is fed into a single-channel anchored deep attractor network to generate separated signals. And the final separation is acquired by post selecting the separating output for each beams. To evaluate the proposed system, we create a challenging dataset comprising mixtures of 2, 3 or 4 speakers. Our results show that the proposed system largely improves the state of the art in speech separation, achieving 11.5 dB, 11.76 dB and 11.02 dB average signal-to-distortion ratio improvement for 4, 3 and 2 overlapped speaker mixtures, which is comparable to the performance of a minimum variance distortionless response beamformer that uses oracle location, source, and noise information. We also run speech recognition with a clean trained acoustic model on the separated speech, achieving relative word error rate (WER) reduction of 45.76\%, 59.40\% and 62.80\% on fully overlapped speech of 4, 3 and 2 speakers, respectively. With a far talk acoustic model, the WER is further reduced.

研究动机与目标

  • 为解决单通道语音分离在复杂、高度重叠的多说话人场景(尤其是4人及以上)中的局限性。
  • 结合空间信息(波束成形)与频谱信息(深度吸引子)以提升分离性能。
  • 开发一种可扩展、端到端兼容的系统,能够处理可变数量的说话人,并对混响和噪声具有鲁棒性。
  • 证明多通道波束成形后接单通道分离的方法可达到或超越使用理想空间信息的系统性能。

提出的方法

  • 将多通道语音混合输入通过固定差分波束成形器转换为一组波束成形信号,该方法对空间和频谱失配更具鲁棒性。
  • 每个波束成形信号由独立的单通道锚定深度吸引子网络(DAN)处理,该网络学习说话人特定的嵌入向量以实现源分离。
  • 最终分离输出通过后处理方式选择,即为每个说话人选取SIR最高的波束,避免排列模糊问题。
  • 系统通过联合使用嵌入向量的监督损失和基于SIR的波束选择进行训练,实现对分离质量的端到端优化。
  • 构建了一个新型数据集,包含2、3和4人混合语音,用于在真实且具有挑战性的条件下评估性能。

实验结果

研究问题

  • RQ1将多通道波束成形与单通道深度吸引子网络结合,是否能显著提升复杂多说话人语音分离的性能?
  • RQ2与使用理想空间信息的系统相比,所提出的MBDAN系统在SIR和词错误率(WER)方面表现如何?
  • RQ3波束成形步骤在高度重叠的混合语音中,能在多大程度上缓解混响和频谱掩蔽的影响?
  • RQ4所提出的方法是否能在不依赖源位置或干净信号信息的情况下,实现与理想波束成形器相当的性能?

主要发现

  • 所提出的MBDAN系统在4人混合语音中实现11.5 dB的SIR提升,3人混合中为11.76 dB,2人混合中为11.02 dB,显著优于先前的单通道方法。
  • 该系统的性能与使用理想源位置、噪声和语音信息的最小方差无失真响应(MVDR)波束成形器相当。
  • 使用干净语音训练的声学模型进行语音识别,分离后4人混合语音的WER相对降低62.80%,3人混合为59.40%,2人混合为45.76%。
  • 使用远场声学模型后,WER降低进一步提升至69.51%、64.19%和52.53%,归因于对混响和噪声更强的鲁棒性。
  • 最强与最弱说话人之间性能差距为4 dB,表明系统对主导说话人的分离能力极强。
  • 由于波束成形步骤有效降低了频谱掩蔽,该方法显著优于单通道深度吸引子网络,尤其在混响条件下。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。