Skip to main content
QUICK REVIEW

[논문 리뷰] Temporal-Spatial Neural Filter: Direction Informed End-to-End Multi-channel Target Speech Separation

Rongzhi Gu, Yuexian Zou|arXiv (Cornell University)|2020. 01. 02.
Speech and Audio Processing참고 문헌 66인용 수 10
한 줄 요약

이 논문은 반향 환경에서 엔드 투 엔드 다중 채널 목표 음성 분리에 적합한 시간-공간 신경 필터를 제안한다. 시간, 주파수, 공간 및 방향 특징의 통합 모델링을 통해 분리 정확도를 향상시키며, 저지연, 단일 통과 추론을 위한 완전 컨volution 신경망을 사용한다. 방향 추정 오차에 대해 강건하며, 15° 이상 간격을 두고 분리된 경우 기존 최고 수준의 방법들을 능가하며, 파rameter 수가 적고 빠른 속도를 구현한다.

ABSTRACT

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main challenges are the complex acoustic environment and the real-time processing requirement. To address these challenges, we propose a temporal-spatial neural filter, which directly estimates the target speech waveform from multi-speaker mixture in reverberant environments, assisted with directional information of the speaker(s). Firstly, against variations brought by complex environment, the key idea is to increase the acoustic representation completeness through the jointly modeling of temporal, spectral and spatial discriminability between the target and interference source. Specifically, temporal, spectral, spatial along with the designed directional features are integrated to create a joint acoustic representation. Secondly, to reduce the latency, we design a fully-convolutional autoencoder framework, which is purely end-to-end and single-pass. All the feature computation is implemented by the network layers and operations to speed up the separation procedure. Evaluation is conducted on simulated reverberant dataset WSJ0-2mix and WSJ0-3mix under speaker-independent scenario. Experimental results demonstrate that the proposed method outperforms state-of-the-art deep learning based multi-channel approaches with fewer parameters and faster processing speed. Furthermore, the proposed temporal-spatial neural filter can handle mixtures with varying and unknown number of speakers and exhibits persistent performance even when existing a direction estimation error. Codes and models will be released soon.

연구 동기 및 목표

  • 복잡한 음향 조건에서 반향적이고 다중 화자 환경에서의 실제 음성 분리 과제를 해결한다.
  • 시간-주파수 마스킹 방법의 한계를 극복하기 위해 위상 재구성 오류를 방지하기 위해 직접적으로 시간 도메인 웨이브폼을 추정한다.
  • 완전 컨volution 신경망 기반의 단일 통과 엔드 투 엔드 아키텍처를 통해 저지연 실시간 처리를 가능하게 한다.
  • 시간, 주파수, 공간 및 방향 특징의 통합 모델링을 통해 분리 정확도를 향상시킨다.
  • 특히 화자들이 방향에서 가까이 위치하지 않은 경우에 강건한 방향 추정 오차에 대응한다.

제안 방법

  • 시간, 주파수, 공간 및 방향 정보 특징을 통합하여 소스 식별 능력을 향상시킨 통합 음향 표현을 구축한다.
  • 채널 간 위상 차이와 방향 그리드 해상도를 바탕으로 각 시간-주파수 영역에서 소스의 우세도를 나타내는 두 가지 방향 특징(cosIPD 및 DPR)을 설계한다.
  • 입력 혼합 신호를 단일 통과로 처리하는 완전 컨볼루션 오토에인코더 네트워크를 구현하여 실시간 저지연 추론을 가능하게 한다.
  • 중간 T-F 마스크 추정을 생략하고, 다중 채널 혼합 신호로부터 직접 목표 음성 웨이브폼을 예측하도록 모델을 엔드 투 엔드로 훈련한다.
  • 화자 방향을 입력으로 사용하여 네트워크가 목표 소스에 집중하도록 유도하고, 화자별 출력 할당을 가능하게 한다.
  • 방향 특징을 활용해 공간 식별 능력을 향상시키며, 특히 각도 차이가 15°를 초과할 경우에 효과적이다.

실험 결과

연구 질문

  • RQ1시간, 주파수, 공간 및 방향 특징의 통합 모델링이 반향 환경에서 목표 음성 분리 정확도를 향상시킬 수 있는가?
  • RQ2화자들이 밀접하게 분리된 경우, 특히 방향 추정 오차가 발생할 때 제안된 방법의 성능은 어떠한가?
  • RQ3방향 특징의 사용이 화자 위치 추정의 각도 불확실성에 대해 얼마나 강건한가?
  • RQ4완전 컨볼루션 단일 통과 아키텍처가 모델 크기를 줄이고 지연을 낮추면서도 높은 분리 성능을 달성할 수 있는가?
  • RQ5모델은 알려지지 않거나 변동하는 화자 수를 가진 혼합 신호를 어떻게 처리하는가?

주요 결과

  • 제안된 방법은 화자 독립 조건 하에서 WSJ0-2mix 및 WSJ0-3mix 데이터셋에서 최고 수준의 딥러닝 기반 다중 채널 음성 분리 방법들을 능가한다.
  • 방향 추정이 정확하고 화자 각도가 15° 초과일 경우, cosIPD+AF 기준선 대비 SI-SDRi 성능 향상이 0.3 dB 이상이다.
  • 방향 추정 오차로 인한 성능 저하는 미미하다. 각도 차이가 15° 초과이고 오차가 ±10° 이내일 경우, 2-mix에서는 0.1 dB 이하, 3-mix에서는 0.4 dB 이하의 SI-SDRi 저하가 발생한다.
  • 모든 테스트된 방향 추정 오차 조건에서 WSJ0-2mix에서는 1.2 dB 이하, WSJ0-3mix에서는 1.7 dB 이하의 성능 저하를 유지하며 강건성을 확보한다.
  • 기존 엔드 투 엔드 방법들보다 처리 속도가 빠르고 파rameter 수가 적어 실시간 구현에 적합하다.
  • 모델은 최대 세 명의 동시 화자 혼합 신호를 성공적으로 처리하며, 방향 입력에 기반해 출력을 올바른 화자에 할당한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.