[論文レビュー] Epoch-Synchronous Overlap-Add (ESOLA) for Time- and Pitch-Scale Modification of Speech Signals
本稿では、音声信号の時間的・ピッチスケール変更において、計算効率が高く、高品質な手法であるエポック同期オーバーラップアド(ESOLA)を提案する。エポック(声門閉鎖瞬間)にフレームを同期させることで、0.5〜2.0のスケール要因において正確な時間スケーリングを可能にし、PSOLA や WSOLA よりも優れた聴取的品質を達成するとともに、SOLAFS よりも最低3倍速い実行時間を実現する。
Time- and pitch-scale modifications of speech signals find important applications in speech synthesis, playback systems, voice conversion, learning/hearing aids, etc.. There is a requirement for computationally efficient and real-time implementable algorithms. In this paper, we propose a high quality and computationally efficient time- and pitch-scaling methodology based on the glottal closure instants (GCIs) or epochs in speech signals. The proposed algorithm, termed as epoch-synchronous overlap-add time/pitch-scaling (ESOLA-TS/PS), segments speech signals into overlapping short-time frames and then the adjacent frames are aligned with respect to the epochs and the frames are overlap-added to synthesize time-scale modified speech. Pitch scaling is achieved by resampling the time-scaled speech by a desired sampling factor. We also propose a concept of epoch embedding into speech signals, which facilitates the identification and time-stamping of samples corresponding to epochs and using them for time/pitch-scaling to multiple scaling factors whenever desired, thereby contributing to faster and efficient implementation. The results of perceptual evaluation tests reported in this paper indicate the superiority of ESOLA over state-of-the-art techniques. ESOLA significantly outperforms the conventional pitch synchronous overlap-add (PSOLA) techniques in terms of perceptual quality and intelligibility of the modified speech. Unlike the waveform similarity overlap-add (WSOLA) or synchronous overlap-add (SOLA) techniques, the ESOLA technique has the capability to do exact time-scaling of speech with high quality to any desired modification factor within a range of 0.5 to 2. Compared to synchronous overlap-add with fixed synthesis (SOLAFS), the ESOLA is computationally advantageous and at least three times faster.
研究の動機と目的
- 音声処理アプリケーションにおけるリアルタイム性、計算効率、高品質な時間的・ピッチスケール変更のニーズに対応すること。
- ピッチの不一致、音声の聞き取りにくさ、高い計算コストといった、PSOLA や WSOLA などの既存手法の限界を克服すること。
- 声門閉鎖瞬間(エポック)を用いたフレーム同期により、固定の合成フレーム長を維持しながら正確な時間スケーリングを可能にすること。
- エポック埋め込みを導入し、再処理を伴わずにマルチファクタスケーリングをサポートするフレーム同期の高速化を実現すること。
- SOLAFS や WSOLA などの最先端手法と比較して、優れた聴取的品質と計算効率を達成すること。
提案手法
- ピッチに依存しない窓関数を用いて音声を重複する短時間フレームに分割し、重複率をスケール要因で制御する。
- 連続するフレームを声門閉鎖瞬間(GCI)またはエポックに合わせることで、オーバーラップアド合成におけるピッチの一貫性を確保する。
- 固定フレーム長を維持しながら、サンプルの挿入/削除により合成フレームシフトを調整することで、時間スケーリングを実現する。
- 所望のピッチスケーリング要因を用いて、時間スケーリング済み音声信号をリサンプリングすることでピッチスケーリングを達成する。
- エポック情報を音声信号に埋め込むことで、GCI の高速かつ再現可能な特定が可能になり、効率的なフレーム同期が実現する。
- 相互相関や自己相関ではなくエポックに基づく同期を採用することで、計算複雑度を低減し、精度を向上させる。
実験結果
リサーチクエスチョン
- RQ1エポック同期フレーム同期は、PSOLA や WSOLA と比較して、時間的・ピッチスケール変更音声の聴取的品質と話者の理解度を向上させることができるか?
- RQ2WSOLA が期間の一貫性を欠くのに対し、ESOLA は 0.5 から 2.0 の間の任意の要因に対して正確な時間スケーリングを達成できるか?
- RQ3再処理を伴わないマルチファクタスケーリングにおいて、エポック埋め込みが計算コストを顕著に低減するか?
- RQ4SOLAFS や他の最先端技術と比較して、ESOLA の実行時間と品質はどの程度か?
- RQ5ESOLA は、多様なピッチおよび時間スケーリング要因において、最小限のアーティファクトで高品質な出力を維持できるか?
主な発見
- ピッチスケーリング要因 2.0 において、ESOLA は平均評価得点(MOS)4.40 を達成し、LP-PSOLA(1.60)や TD-PSOLA(2.40)を著しく上回った。
- 時間スケーリング要因 1.5 において、ESOLA は MOS 4.65 を達成し、WSOLA(3.25)や SOLAFS(4.10)を上回った。
- SOLAFS と比較して、ESOLA は実行時間を最低3倍短縮し、評価された手法の中で最も高速であった。
- 聴取評価の結果、ESOLA はすべてのスケーリング要因において、PSOLA や WSOLA、SOLAFS よりも高い MOS 値を一貫して達成した。
- エポックベースの同期により、固定フレーム長を維持しながら正確な時間スケーリングが可能になった。これに対して WSOLA は期間の一貫性を欠いた。
- エポック埋め込みにより、GCI の効率的かつ再現可能な特定が可能になり、計算コストの低減と高速なマルチファクタスケーリングが実現された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。