Skip to main content
QUICK REVIEW

[论文解读] Separating Sounds from a Single Image.

Lingyu Zhu, Esa Rahtu|arXiv (Cornell University)|Jul 15, 2020
Speech and Audio Processing参考文献 31被引用 6
一句话总结

本文提出了一种高效的外观注意力模块,利用单张图像的视觉外观来增强声音源分离,从而在不增加额外计算成本的情况下提升语义表征区分度和源定位性能。在MUSIC数据集上,该方法通过有效利用视觉线索和子网络容量分析,实现了最先进性能。

ABSTRACT

Recently, visual information has been widely used to aid the sound source separation tasks. It aims at identifying sound components from a given sound mixture with the presence of visual information. Especially, the appearance cues play an important role on separating sounds. However, the capacity of how well the network processes each modality is often ignored. In this paper, we investigate the performance of appearance information, extracted from a single image, in the task of recovering the original component signals from a mixture audio. An efficient appearance attention module is introduced to improve the sound separation performance by enhancing the distinction of the predicted semantic representations, and to precisely locate sound sources without extra computation. Moreover, we utilize the ground category information to study the capacity of each sub-network. We compare the proposed methods with recent baselines on the MUSIC dataset. Project page: this https URL

研究动机与目标

  • 研究单张图像的视觉外观在多大程度上能提升声音源分离性能。
  • 设计一种高效的注意力机制,以提升语义表征区分度和源定位性能。
  • 利用真实类别信息分析各个子网络的容量。
  • 在MUSIC数据集上与近期基线方法进行比较,以验证性能。

提出的方法

  • 引入外观注意力模块,以优化视觉特征并增强声音分离中的语义表征区分度。
  • 该模块通过利用单张图像中已有的视觉特征,无需额外计算即可运行。
  • 使用真实类别信息来评估和分析模型中各子网络的容量。
  • 该方法在MUSIC数据集上进行训练和评估,该数据集提供了包含多个声音源的视听混合数据。
  • 通过注意力机制聚焦于与声音源相关的判别性视觉区域,实现特征增强。
  • 该架构通过共享注意力机制整合视觉与音频模态表征,以提升分离准确率。

实验结果

研究问题

  • RQ1单张图像的外观线索在多大程度上能有效提升声音源分离性能?
  • RQ2注意力机制是否能在不增加计算成本的情况下增强语义表征区分度?
  • RQ3各子网络的容量在多大程度上影响整体分离性能?
  • RQ4真实类别信息对子网络容量分析有何影响?

主要发现

  • 所提出的外观注意力模块在MUSIC数据集上显著提升了声音源分离性能。
  • 该方法通过利用视觉线索增强语义表征区分度,实现了最先进结果。
  • 注意力机制在不增加计算成本的前提下实现了精确的源定位。
  • 基于真实类别信息的子网络容量分析揭示了各组件之间有意义的性能差异。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。