Skip to main content
QUICK REVIEW

[论文解读] Classification with Joint Time-Frequency Scattering.

Joakim Andén, Vincent Lostanlen|arXiv (Cornell University)|Jul 24, 2018
Music and Audio Processing参考文献 50被引用 9
一句话总结

本文提出了联合时频散射变换,一种具有固定小波基滤波器的深度卷积神经网络,能够捕捉时间与频率域中的多尺度能量分布。该方法在音频分类任务中实现了最先进性能,优于时域散射方法,并在准确率上与可学习网络相当。

ABSTRACT

In time series classification, signals are typically mapped into some intermediate representation which is used to construct models. We introduce the joint time-frequency scattering transform, a locally time-shift invariant representation which characterizes the multiscale energy distribution of a signal in time and frequency. It is computed through wavelet convolutions and modulus non-linearities and may therefore be implemented as a deep convolutional neural network whose filters are not learned but calculated from wavelets. We consider the progression from mel-spectrograms to time scattering and joint time-frequency scattering transforms, illustrating the relationship between increased discriminability and refinements of convolutional network architectures. The suitability of the joint time-frequency scattering transform for characterizing time series is demonstrated through applications to chirp signals and audio synthesis experiments. The proposed transform also obtains state-of-the-art results on several audio classification tasks, outperforming time scattering transforms and achieving accuracies comparable to those of fully learned networks.

研究动机与目标

  • 开发一种时间序列的时频不变表示,以增强分类任务中的判别能力。
  • 通过引入结构化、非学习型特征提取器,弥合传统谱图与可学习深度网络之间的差距。
  • 证明联合时频散射相较于仅时域散射的优越性,并实现与完全训练网络相当的性能。
  • 分析从梅尔谱图到时域散射,再到联合时频散射的架构演进过程。

提出的方法

  • 联合时频散射变换通过小波卷积和模非线性运算计算,生成局部时移不变表示。
  • 滤波器基于小波而非学习获得,实现固定且可解释的特征提取流程。
  • 架构设计为无可训练参数的深度卷积神经网络,依赖预定义的小波滤波器。
  • 该方法捕捉信号在时间和频率域中的多尺度能量分布。
  • 在时域散射基础上扩展,引入频率偏移不变性,提升鲁棒性与判别能力。

实验结果

研究问题

  • RQ1在音频分类中,联合时频散射相较于仅时域散射如何提升判别能力?
  • RQ2固定且非学习的小波基网络在多大程度上可实现与完全训练深度网络相当的性能?
  • RQ3从梅尔谱图到时域散射,再到联合时频散射的演进过程,如何影响特征质量与分类准确率?
  • RQ4联合时频不变性在增强对信号扰动的鲁棒性方面起到何种作用?

主要发现

  • 联合时频散射变换在多个音频分类基准上实现了最先进性能。
  • 其性能优于时域散射变换,证明了引入频域不变性的优势。
  • 尽管无可训练参数,该方法仍实现了与完全学习的深度神经网络相当的分类准确率。
  • 该变换能有效表征 chirp 信号与音频合成数据,展现出强大的表征能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。