Skip to main content
QUICK REVIEW

[论文解读] Epoch-Synchronous Overlap-Add (ESOLA) for Time- and Pitch-Scale Modification of Speech Signals

Sunil Rudresh, Aditya Vasisht|arXiv (Cornell University)|Jan 19, 2018
Speech and Audio Processing参考文献 20被引用 14
一句话总结

本文提出了一种计算效率高且音质优良的时间-音高比例变换方法——时钟同步重叠相加(ESOLA)。通过将帧同步到声门闭合瞬时(时钟点),ESOLA 实现了 0.5 至 2.0 之间的精确时间缩放,其主观音质优于 PSOLA 和 WSOLA,且执行时间相比 SOLAFS 至少减少三倍。

ABSTRACT

Time- and pitch-scale modifications of speech signals find important applications in speech synthesis, playback systems, voice conversion, learning/hearing aids, etc.. There is a requirement for computationally efficient and real-time implementable algorithms. In this paper, we propose a high quality and computationally efficient time- and pitch-scaling methodology based on the glottal closure instants (GCIs) or epochs in speech signals. The proposed algorithm, termed as epoch-synchronous overlap-add time/pitch-scaling (ESOLA-TS/PS), segments speech signals into overlapping short-time frames and then the adjacent frames are aligned with respect to the epochs and the frames are overlap-added to synthesize time-scale modified speech. Pitch scaling is achieved by resampling the time-scaled speech by a desired sampling factor. We also propose a concept of epoch embedding into speech signals, which facilitates the identification and time-stamping of samples corresponding to epochs and using them for time/pitch-scaling to multiple scaling factors whenever desired, thereby contributing to faster and efficient implementation. The results of perceptual evaluation tests reported in this paper indicate the superiority of ESOLA over state-of-the-art techniques. ESOLA significantly outperforms the conventional pitch synchronous overlap-add (PSOLA) techniques in terms of perceptual quality and intelligibility of the modified speech. Unlike the waveform similarity overlap-add (WSOLA) or synchronous overlap-add (SOLA) techniques, the ESOLA technique has the capability to do exact time-scaling of speech with high quality to any desired modification factor within a range of 0.5 to 2. Compared to synchronous overlap-add with fixed synthesis (SOLAFS), the ESOLA is computationally advantageous and at least three times faster.

研究动机与目标

  • 解决语音处理应用中对实时、计算高效且音质优良的时间-音高比例变换的需求。
  • 克服现有技术(如 PSOLA 和 WSOLA)存在的音高不一致、可懂度差和计算成本高等局限。
  • 通过利用声门闭合瞬时(时钟点)进行帧对齐,实现固定合成帧长下的精确时间缩放。
  • 引入时钟嵌入技术以加速帧对齐,并支持无需重新处理的多因子缩放。
  • 与最先进的方法(如 SOLAFS 和 WSOLA)相比,实现更优的主观音质和计算效率。

提出的方法

  • 使用与音高无关的窗函数将语音分割为重叠的短时帧,重叠程度由时间缩放因子控制。
  • 基于声门闭合瞬时(GCI)或时钟点对连续帧进行对齐,以确保重叠相加合成过程中的音高一致性。
  • 通过调整合成帧移量(保持帧长固定)实现时间缩放,方法为插入/删除样本。
  • 通过使用所需音高缩放因子对时间缩放后的语音信号进行重采样,实现音高缩放。
  • 将时钟信息嵌入语音信号中,以实现快速、可重复的 GCI 识别,从而提升帧对齐效率。
  • 采用基于时钟的对齐方式,而非互相关或自相关,降低计算复杂度并提高准确性。

实验结果

研究问题

  • RQ1与 PSOLA 和 WSOLA 相比,基于时钟同步的帧对齐是否能提升时间-音高缩放语音的主观音质和可懂度?
  • RQ2ESOLA 是否能实现 0.5 至 2.0 之间任意因子的精确时间缩放,而克服 WSOLA 缺乏持续时间一致性的缺陷?
  • RQ3与重新处理相比,时钟嵌入是否能显著降低多因子缩放的计算成本?
  • RQ4与 SOLAFS 及其他最先进方法相比,ESOLA 在执行时间和音质方面表现如何?
  • RQ5ESOLA 是否能在各种音高和时间缩放因子下保持高质量输出,且引入极少伪影?

主要发现

  • 在音高缩放因子为 2.0 时,ESOLA 的平均意见得分(MOS)达到 4.40,显著优于 LP-PSOLA(1.60)和 TD-PSOLA(2.40)。
  • 在时间缩放因子为 1.5 时,ESOLA 的 MOS 达到 4.65,优于 WSOLA(3.25)和 SOLAFS(4.10)。
  • 与 SOLAFS 相比,ESOLA 将执行时间至少减少三倍,是所评估方法中最快的一种。
  • 主观评价显示,ESOLA 在所有缩放因子下均持续提供高于 PSOLA、WSOLA 和 SOLAFS 的 MOS 值。
  • 基于时钟的对齐方式实现了固定帧长下的精确时间缩放,而 WSOLA 则缺乏持续时间一致性。
  • 时钟嵌入使得 GCI 的识别更加高效且可重复,降低了计算成本,并支持快速的多因子缩放。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。