Skip to main content
QUICK REVIEW

[论文解读] Structured Sparsity Models for Multiparty Speech Recovery from Reverberant Recordings

Afsaneh Asaei, Mohammad Golbabaee|arXiv (Cornell University)|Oct 25, 2012
Speech and Audio Processing参考文献 35被引用 5
一句话总结

该论文通过利用空间、频谱和声学结构,提出了一种用于混响环境中多人语音恢复的结构化稀疏模型。通过虚拟源像的块稀疏恢复估计房间几何结构和吸声系数,结合结构化稀疏逼近与逆滤波,实现高精度语音分离与识别。真实数据实验表明,在三名说话人条件下,词识别率最高可达79.2%。

ABSTRACT

We tackle the multi-party speech recovery problem through modeling the acoustic of the reverberant chambers. Our approach exploits structured sparsity models to perform room modeling and speech recovery. We propose a scheme for characterizing the room acoustic from the unknown competing speech sources relying on localization of the early images of the speakers by sparse approximation of the spatial spectra of the virtual sources in a free-space model. The images are then clustered exploiting the low-rank structure of the spectro-temporal components belonging to each source. This enables us to identify the early support of the room impulse response function and its unique map to the room geometry. To further tackle the ambiguity of the reflection ratios, we propose a novel formulation of the reverberation model and estimate the absorption coefficients through a convex optimization exploiting joint sparsity model formulated upon spatio-spectral sparsity of concurrent speech representation. The acoustic parameters are then incorporated for separating individual speech signals through either structured sparse recovery or inverse filtering the acoustic channels. The experiments conducted on real data recordings demonstrate the effectiveness of the proposed approach for multi-party speech recovery and recognition.

研究动机与目标

  • 解决在直达路径信号被早期反射遮蔽的多人混响录音中恢复个体语音信号的挑战。
  • 通过在自由场模型中对空间谱进行稀疏逼近,实现虚拟源的定位,从而建模房间声学特性。
  • 利用时频分量中的低秩结构对每位说话人的图像进行聚类,以推断源特定子空间。
  • 通过跨频带联合稀疏性的凸优化方法估计吸声系数。
  • 通过将估计的声学参数整合到结构化稀疏恢复或逆滤波中,实现鲁棒的语音分离与识别。

提出的方法

  • 将房间离散化为网格单元,将声源位置建模为仅含N个非零条目的稀疏向量(N为说话人数量)。
  • 基于多通道录音的空间谱,利用稀疏逼近在自由场模型中定位虚拟源。
  • 利用时频分量中低秩结构对每位说话人的早期图像进行聚类,以识别源特定子空间。
  • 在频带间建立联合稀疏模型,通过凸优化估计吸声系数。
  • 利用估计的房间脉冲响应进行结构化稀疏恢复或逆滤波,以重构个体语音信号。
  • 在谱系数中引入谐波特性与时间相关性,以提升恢复性能。

实验结果

研究问题

  • RQ1通过利用空间与频谱结构,结构化稀疏模型能否有效恢复多人混响环境中的语音信号?
  • RQ2如何通过稀疏逼近方法,基于早期虚拟源像的定位推断房间几何结构?
  • RQ3在混响房间中,跨频带的联合稀疏性在多大程度上改善了吸声系数的估计?
  • RQ4在多说话人场景中,结构化稀疏恢复或逆滤波是否优于传统波束成形?
  • RQ5如谐波特性与低秩时频分量等参数化结构,在多大程度上提升了语音恢复性能?

主要发现

  • 所提出的RAM-SR方法在三名说话人同时说话条件下,实现了79.21%的词识别率,显著优于基线波束成形(39.92%)和RIR-LS方法(70.88%)。
  • 在三名说话人场景中,RAM-SR的信干比(SIR)达到14.2 dB,而RIR-LS为10.1 dB,基线方法为-0.7 dB。
  • 在三名说话人条件下,感知语音质量评价(PESQ)得分提升至2.62(RAM-SR),而基线方法仅为1.6。
  • 该方法成功利用MONC语料库中9,000段语音,以4 Hz的分辨率估计出每面墙的频率相关吸声系数。
  • 块稀疏恢复与低秩聚类的结合,实现了准确的房间几何结构推断,并提升了语音分离性能。
  • 实验结果证实,结合空间、频谱与声学约束的结构化稀疏模型,显著增强了真实混响录音中的语音恢复性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。