Skip to main content
QUICK REVIEW

[论文解读] Sparse, complex-valued representations of natural sounds learned with phase and amplitude continuity priors

Wiktor Młynarski|arXiv (Cornell University)|Dec 17, 2013
Hearing Loss and Rehabilitation参考文献 27被引用 4
一句话总结

本文通过在幅度和相位上施加时间连续性先验,提出了一种相位不变的复值稀疏编码方法,用于自然声音。该方法能够实现高效、鲁棒的表示,并将相位作为时间延迟显式保留。尽管自然声音本身不具备固有的相位不变性,但该方法学习到的字典在去噪性能上与无约束模型相当,同时显式编码了相位信息,使其适用于需要相位敏感性的感知与机器学习任务。

ABSTRACT

Complex-valued sparse coding is a data representation which employs a dictionary of two-dimensional subspaces, while imposing a sparse, factorial prior on complex amplitudes. When trained on a dataset of natural image patches, it learns phase invariant features which closely resemble receptive fields of complex cells in the visual cortex. Features trained on natural sounds however, rarely reveal phase invariance and capture other aspects of the data. This observation is a starting point of the present work. As its first contribution, it provides an analysis of natural sound statistics by means of learning sparse, complex representations of short speech intervals. Secondly, it proposes priors over the basis function set, which bias them towards phase-invariant solutions. In this way, a dictionary of complex basis functions can be learned from the data statistics, while preserving the phase invariance property. Finally, representations trained on speech sounds with and without priors are compared. Prior-based basis functions reveal performance comparable to unconstrained sparse coding, while explicitely representing phase as a temporal shift. Such representations can find applications in many perceptual and machine learning tasks.

研究动机与目标

  • 使用稀疏复值表征分析自然声音的高阶统计特性。
  • 开发促进自然声音学习字典中相位不变性的先验。
  • 在听觉数据表征中显式表示相位作为时间延迟。
  • 评估相位不变字典在编码效率和去噪性能方面是否可与无约束模型相匹配。

提出的方法

  • 在复数域中构建基于复数基函数的复值稀疏编码,将信号表示为复数幅度与基函数的线性组合。
  • 通过惩罚基函数在相位和幅度上的时间变化性,引入相位与幅度的连续性先验,促进其随时间缓慢、平滑地变化。
  • 使用改进的K-SVD算法在这些先验下学习过完备和完备字典,确保相位不变性的同时保持数据保真度。
  • 采用两阶段优化:首先在先验约束下学习基函数,然后通过阈值法计算稀疏系数。
  • 通过系数熵和在语音语料上的去噪性能评估编码效率,比较基于先验与无约束模型的性能。
  • 使用归一化直方图估计Kullback-Leibler散度与熵,评估模型输出分布与真实数据分布的接近程度。

实验结果

研究问题

  • RQ1自然声音的高阶统计特性与自然图像在相位不变性方面有何不同?
  • RQ2尽管自然声音本身不具有内在的相位不变性,能否通过结构化先验从自然声音数据中学习到相位不变特征?
  • RQ3对基函数的幅度和相位施加时间连续性约束,如何影响学习到的字典的结构与性能?
  • RQ4基于先验的字典在多大程度上可匹配无约束模型的编码效率与去噪性能?

主要发现

  • 与图像数据不同,相位不变特征在自然声音中并非自然学习到的,这是由于音频中存在强烈的跨频带相关性。
  • 所提出的幅度与相位连续性先验成功在学习到的基函数中诱导出相位不变性,即使自然声音本身并不天然倾向于此类不变性。
  • 基于先验的字典在去噪性能上与无约束模型相当,表明表征质量未受损。
  • 过完备表征的系数熵低于完备表征,表明其编码效率更高,这与图像模型中的发现相反。
  • 尽管性能相近,无约束模型的熵略低,表明其可能更接近真实的数据生成分布。
  • 该方法能够显式表示相位作为时间延迟,这对空间听觉与声源定位等任务至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。