Skip to main content
QUICK REVIEW

[논문 리뷰] Epoch-Synchronous Overlap-Add (ESOLA) for Time- and Pitch-Scale Modification of Speech Signals

Sunil Rudresh, Aditya Vasisht|arXiv (Cornell University)|2018. 01. 19.
Speech and Audio Processing참고 문헌 20인용 수 14
한 줄 요약

이 논문은 음성 신호의 시간 및 편도 스케일 변환을 위한 계산적으로 효율적이고 고품질인 방법인 에포크 동기 오버랩-애드(ESOLA)를 제안한다. 갑상선 폐쇄 순간(에포크)에 프레임을 동기화시킴으로써, 0.5에서 2.0 사이의 인자로 정확한 시간 스케일링이 가능하며, PSOLA 및 WSOLA보다 청각적 품질이 뛰어나고, SOLAFS보다 실행 시간을 최소 3배 이상 단축시킨다.

ABSTRACT

Time- and pitch-scale modifications of speech signals find important applications in speech synthesis, playback systems, voice conversion, learning/hearing aids, etc.. There is a requirement for computationally efficient and real-time implementable algorithms. In this paper, we propose a high quality and computationally efficient time- and pitch-scaling methodology based on the glottal closure instants (GCIs) or epochs in speech signals. The proposed algorithm, termed as epoch-synchronous overlap-add time/pitch-scaling (ESOLA-TS/PS), segments speech signals into overlapping short-time frames and then the adjacent frames are aligned with respect to the epochs and the frames are overlap-added to synthesize time-scale modified speech. Pitch scaling is achieved by resampling the time-scaled speech by a desired sampling factor. We also propose a concept of epoch embedding into speech signals, which facilitates the identification and time-stamping of samples corresponding to epochs and using them for time/pitch-scaling to multiple scaling factors whenever desired, thereby contributing to faster and efficient implementation. The results of perceptual evaluation tests reported in this paper indicate the superiority of ESOLA over state-of-the-art techniques. ESOLA significantly outperforms the conventional pitch synchronous overlap-add (PSOLA) techniques in terms of perceptual quality and intelligibility of the modified speech. Unlike the waveform similarity overlap-add (WSOLA) or synchronous overlap-add (SOLA) techniques, the ESOLA technique has the capability to do exact time-scaling of speech with high quality to any desired modification factor within a range of 0.5 to 2. Compared to synchronous overlap-add with fixed synthesis (SOLAFS), the ESOLA is computationally advantageous and at least three times faster.

연구 동기 및 목표

  • 음성 처리 응용 분야에서 실시간, 계산적으로 효율적이며 고품질의 시간 및 편도 스케일 변환을 위한 필요를 해결한다.
  • 편도 일관성 결여, 낮은 이해도, 높은 계산 비용 등의 문제를 겪는 기존 기법들인 PSOLA 및 WSOLA의 한계를 극복한다.
  • 갑상선 폐쇄 순간(에포크)을 활용해, 고정된 합성 프레임 길이로 정확한 시간 스케일링을 가능하게 한다.
  • 에포크 임bed딩을 도입하여, 재처리 없이도 다중 인자 스케일링을 지원하는 빠른 프레임 정렬을 가능하게 한다.
  • SOLAFS 및 WSOLA와 같은 최첨단 기법들과 비교해 뛰어난 청각적 품질과 계산 효율성을 달성한다.

제안 방법

  • 시간 스케일 인자에 따라 제어되는 오버랩을 갖는 피치 무관 윈도잉을 사용해 음성을 겹치는 짧은 시간 프레임으로 분할한다.
  • 연속된 프레임을 갑상선 폐쇄 순간(GCIs) 또는 에포크에 기반해 정렬하여 오버랩-애드 합성 중 편도 일관성을 확보한다.
  • 고정된 프레임 길이를 유지하면서 샘플 삽입/삭제를 통해 합성 프레임 이격을 조정함으로써 시간 스케일 변환을 수행한다.
  • 원하는 편도 스케일 인자로 시간 스케일링된 음성 신호를 재샘플링하여 편도 스케일링을 달성한다.
  • 에포크 정보를 음성 신호에 임베딩하여 GCIs의 신속하고 반복 가능한 식별을 가능하게 하여 효율적인 프레임 정렬을 지원한다.
  • 교차상관 또는 자기상관을 대체로 에포크 기반 정렬을 사용함으로써 계산 복잡도를 감소시키고 정확도를 향상시킨다.

실험 결과

연구 질문

  • RQ1에포크 기반 동기 프레임 정렬은 PSOLA 및 WSOLA에 비해 시간 및 편도 스케일링된 음성의 청각적 품질과 이해도를 향상시키는가?
  • RQ2ESOLA는 WSOLA가 지속성 결여로 인해 시간 스케일링이 불가능한 0.5에서 2.0 사이의 모든 인자에 대해 정확한 시간 스케일링을 달성할 수 있는가?
  • RQ3에포크 임베딩은 재처리 없이도 다중 인자 스케일링에 대해 계산 비용을 크게 감소시키는가?
  • RQ4ESOLA는 SOLAFS 및 기타 최첨단 기법들과 비교해 실행 시간과 품질 면에서 어떻게 비교되는가?
  • RQ5ESOLA는 다양한 편도 및 시간 스케일 인자에서 최소한의 잡음으로 고품질 출력을 유지할 수 있는가?

주요 결과

  • 편도 스케일 인자 2.0에서 ESOLA는 평균 관점 점수(MOS) 4.40을 기록하여 LP-PSOLA(1.60) 및 TD-PSOLA(2.40)를 크게 앞서며 뛰어난 성능을 보였다.
  • 시간 스케일 인자 1.5에서 ESOLA는 MOS 4.65를 기록하여 WSOLA(3.25) 및 SOLAFS(4.10)를 초월했다.
  • ESOLA는 SOLAFS 대비 실행 시간을 최소 3배 이상 단축시켜 평가된 방법들 중 가장 빠른 속도를 기록했다.
  • 청각 평가 결과, 모든 스케일 인자에서 ESOLA는 PSOLA, WSOLA, SOLAFS보다 일관되게 높은 MOS 값을 제공했다.
  • 에포크 기반 정렬은 고정된 프레임 길이로 정확한 시간 스케일링을 가능하게 하며, WSOLA와 달리 지속성 결여 문제를 해결한다.
  • 에포크 임베딩은 GCIs의 효율적이고 반복 가능한 식별을 가능하게 하여 계산 비용을 감소시키고, 신속한 다중 인자 스케일링을 지원한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.