Skip to main content
QUICK REVIEW

[논문 리뷰] Effective Low-Cost Time-Domain Audio Separation Using Globally Attentive Locally Recurrent Networks

Max W. Y. Lam, Jun Wang|arXiv (Cornell University)|2021. 01. 13.
Speech and Audio Processing참고 문헌 38인용 수 9
한 줄 요약

이 논문은 시간 도메인 음성 분리에 적합한 저비용·고효율 아키텍처인 글로벌 주의 집중형 국소 순환(GALR) 네트워크를 제안한다. 이는 국소 순환 처리와 글로벌 자기주목(self-attention)을 결합하여 장거리 의존성을 모델링한다. GALR는 DPRNN 대비 36.1% 적은 런타임 메모리와 49.4% 적은 FLOPs를 사용하면서도, WSJ0-2mix에서 2.4 dB의 SI-SNRi 향상으로 최신 기준 성능을 달성한다.

ABSTRACT

Recent research on the time-domain audio separation networks (TasNets) has brought great success to speech separation. Nevertheless, conventional TasNets struggle to satisfy the memory and latency constraints in industrial applications. In this regard, we design a low-cost high-performance architecture, namely, globally attentive locally recurrent (GALR) network. Alike the dual-path RNN (DPRNN), we first split a feature sequence into 2D segments and then process the sequence along both the intra- and inter-segment dimensions. Our main innovation lies in that, on top of features recurrently processed along the inter-segment dimensions, GALR applies a self-attention mechanism to the sequence along the inter-segment dimension, which aggregates context-aware information and also enables parallelization. Our experiments suggest that GALR is a notably more effective network than the prior work. On one hand, with only 1.5M parameters, it has achieved comparable separation performance at a much lower cost with 36.1% less runtime memory and 49.4% fewer computational operations, relative to the DPRNN. On the other hand, in a comparable model size with DPRNN, GALR has consistently outperformed DPRNN in three datasets, in particular, with a substantial margin of 2.4dB absolute improvement of SI-SNRi in the benchmark WSJ0-2mix task.

연구 동기 및 목표

  • 산업 적용에서 기존 시간 도메인 음성 분리 모델의 높은 메모리 및 지연 비용 문제를 해결하기 위해.
  • TasNets에서 장시계열 모델링을 향상시키기 위해 RNN과 자기주목 메커니즘의 장점을 융합하기 위해.
  • 저비용 계산 및 메모리 오버헤드를 유지하면서도 높은 분리 성능을 달성하는 컴act하고 효율적인 아키텍처를 설계하기 위해.
  • 자기주목이 음성 시계열에서 글로벌 장거리 의존성을 모델링하는 데 RNN보다 우수한지 검증하기 위해.
  • 다양한 데이터셋에서 일관된 성능 향상을 입증하기 위해, 음성-음성 및 음성-음악 분리 작업을 포함한다.

제안 방법

  • GALR 아키텍처는 입력 웨이브포맷을 2차원 세그먼트로 분할하고, 세그먼트 내 차원을 따라 국소 순환 처리를 적용하여 단기 의존성을 모델링한다.
  • 세그먼트 간 차원을 따라 순환 네트워크를 적용하여 세그먼트 간 순차적 맥락을 포착한다.
  • 세그먼트 간 순서에 대해 자기주목 메커니즘을 적용하여 병렬로 맥락 인식 장거리 의존성을 집계한다.
  • 모델은 원시 웨이브포맷에 직접 시간 도메인 손실을 최적화하기 위해 순열 불변 훈련(PIT)을 활용한다.
  • 기존 모델인 DPRNN과 비교해 FLOPs와 메모리 사용량이 감소한 파rameter 효율적인 아키텍처로 설계되어 있다.
  • 전역 주의 집중의 병렬 처리를 유지하면서도 국소 순환을 보존함으로써 모델링 능력과 효율성의 균형을 이룬다.
Fig. 1 : Upper: a 4s raw waveform mixture of two overlapping utterances; Lower: zooming in on the 385th segment around the lateral phoneme of /l/.
Fig. 1 : Upper: a 4s raw waveform mixture of two overlapping utterances; Lower: zooming in on the 385th segment around the lateral phoneme of /l/.

실험 결과

연구 질문

  • RQ1표준 자기주목의 제곱 복잡도 문제를 해결하면서, 자기주목이 장시간 도메인 음성 시계열에 효과적으로 적용될 수 있는가?
  • RQ2국소 순환과 전역 자기주목을 융합하면 RNN 전용 접근법보다 장거리 의존성을 더 잘 모델링할 수 있는가?
  • RQ3작은 모델로도 최신 기준 모델인 DPRNN 대비 상당히 감소된 FLOPs와 메모리 사용량을 달성하면서도 뛰어난 성능을 낼 수 있는가?
  • RQ4제안된 아키텍처는 음성-음성 및 음성-음악 혼합 분리 작업을 포함한 다양한 분리 작업에서 어떻게 성능을 발휘하는가?
  • RQ5자기주목 메커니즘이 음성 신호에서 먼 맥락 정보를 포착하는 데 RNN보다 더 효과적인가?

주요 결과

  • GALR는 파arameter가 더 적은 상황에서 DPRNN 대비 WSJ0-2mix 데이터셋에서 SI-SNRi가 2.4 dB 향상되었다.
  • 150만 파라미터로만 구성된 GALR는 DPRNN 대비 FLOPs를 49.4% 감소시키고 런타임 메모리를 36.1% 감소시켰지만, 유사한 성능을 유지했다.
  • 더 큰 Libri-2mix 데이터셋에서 GALR는 모델 크기가 더 작았음에도 불구하고(230만 대비 260만 파라미터), SI-SNRi와 SDRi가 각각 0.2 dB 높았다.
  • 음성-음악 혼합 분리 작업에서는 GALR가 DPRNN 대비 SI-SNRi와 SDRi에서 각각 1.4 dB 향상되어 도전적인 상황에서도 뛰어난 강건성을 보였다.
  • 세 데이터셋 전반에서 GALR는 항상 DPRNN을 능가했으며, 이는 전역 자기주목이 장거리 시계열 모델링에 효과적임을 확인시켰다.
  • 결과는 자기주목이 특히 먼 의존성이 중요한 경우, RNN보다 음성 시계열의 글로벌 맥락을 모델링하는 데 더 적합함을 시사한다.
Fig. 2 : Left: the overall architecture of our GALR network. Right: detailed illustration about how the intra- and inter-segment sequences are processed in the locally recurrent layer (lower right) and the globally attentive layer (upper right) inside each GALR block, respectively.
Fig. 2 : Left: the overall architecture of our GALR network. Right: detailed illustration about how the intra- and inter-segment sequences are processed in the locally recurrent layer (lower right) and the globally attentive layer (upper right) inside each GALR block, respectively.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.