[论文解读] An efficient and perceptually motivated auditory neural encoding and decoding algorithm for spiking neural networks
本文提出了一种生物可实现的听觉编码(BAE)方案,通过模拟人类听觉感知,将语音转换为稀疏、高效的脉冲模式,用于脉冲神经网络(SNNs)。通过整合耳蜗滤波、听觉掩蔽效应以及受人类生理启发的脉冲编码,BAE 实现了高保真度的语音重建,并在基于 SNN 的语音识别中表现出稳健性能,从而创建了两个公开的脉冲基数据集:Spike-TIDIGITS 和 Spike-TIMIT。
Auditory front-end is an integral part of a spiking neural network (SNN) when performing auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reconstructable spike pattern to facilitate the subsequent processing. However, most of the auditory front-ends in current studies have not made use of recent findings in psychoacoustics and physiology concerning human listening. In this paper, we propose a neural encoding and decoding scheme that is optimized for speech processing. The neural encoding scheme, that we call Biologically plausible Auditory Encoding (BAE), emulates the functions of the perceptual components of the human auditory system, that include the cochlear filter bank, the inner hair cells, auditory masking effects from psychoacoustic models, and the spike neural encoding by the auditory nerve. We evaluate the perceptual quality of the BAE scheme using PESQ; the performance of the BAE based on speech recognition experiments. Finally, we also built and published two spike-version of speech datasets: the Spike-TIDIGITS and the Spike-TIMIT, for researchers to use and benchmarking of future SNN research.
研究动机与目标
- 开发一种适用于 SNN 的听觉前端,利用近期在人类心理声学与听觉生理学中的发现。
- 通过从语音信号中生成稀疏、高效且可重构的脉冲模式,减轻 SNN 的计算负担。
- 通过采用更具生物真实性的编码方案,提升 SNN 在语音识别任务中的性能。
- 生成并发布标准化的脉冲基语音数据集——Spike-TIDIGITS 和 Spike-TIMIT——以供未来 SNN 研究的基准测试。
提出的方法
- BAE 方案使用伽马函数滤波器组模拟耳蜗频率滤波,以建模人类听觉系统。
- 基于 MPEG-1 Layer III 标准推导的心理声学听觉掩蔽效应被整合,以抑制感知上不相关的成分。
- 通过 15 个均匀分布的阈值上的阈值穿越事件实现脉冲编码,生成时间域脉冲模式。
- 该方法整合了内毛细胞转导与听觉神经脉冲编码,以生成率编码的脉冲输出。
- 该方法旨在保留感知相关的动态特性,同时最小化冗余脉冲,从而提升计算效率。
- 生成的脉冲模式被用于在语音识别任务中训练和评估 SNN。
实验结果
研究问题
- RQ1基于生物启发的听觉编码方案是否能提升 SNN 语音处理的效率与效果?
- RQ2在脉冲编码中引入听觉掩蔽在多大程度上能保持感知质量与识别性能?
- RQ3与传统表示相比,BAE 生成的脉冲模式在可重构性与 SNN 分类准确率方面表现如何?
- RQ4所提出的编码方法是否能在连续语音数据集(如 TIMIT)上实现高性能的 SNN?
- RQ5脉冲模式的稀疏性与时间结构对 SNN 学习与推理有何影响?
主要发现
- BAE 编码实现了高感知质量,PESQ 评分表明语音重建保真度接近透明。
- 在 BAE 编码数据上训练的 SNN 在语音识别任务中表现出色,证明了该编码方案的有效性。
- 听觉掩蔽显著减少了脉冲数量,而未损害神经元响应动态,表现为存在或不存在掩蔽时膜电位轨迹与脉冲输出相似。
- BAE 生成的脉冲模式即使在脉冲率降低的情况下,仍保留了正确分类所必需的关键时间与频谱特征。
- 所提出的 Spike-TIDIGITS 和 Spike-TIMIT 数据集已成功创建并作为公开基准发布,供 SNN 研究使用。
- 结果表明,BAE 实现了高效、基于感知动机的编码,降低了计算负载,同时保持了识别准确率。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。