Skip to main content
QUICK REVIEW

[논문 리뷰] Predict-and-Update Network: Audio-Visual Speech Recognition Inspired by Human Speech Perception

Jiadong Wang, Xinyuan Qian|arXiv (Cornell University)|2022. 09. 05.
Speech and Audio Processing인용 수 6
한 줄 요약

이 논문은 인간의 언어 인지 방식을 영감으로 삼아 시각 신호가 청각 처리를 예측하고 이끌 수 있도록 설계된 종단 간 음성-시각 음성 인식 모델인 Predict-and-Update 네트워크(P&U net)를 제안한다. 크로스모달 Conformer를 통해 시각 임베딩이 청각 표현을 조절함으로써, 이 모델은 기존 방법에 비해 청청 환경에서 10퍼센트 이상, 노이즈가 있는 조건에서 40퍼센트 이상의 WER 감소를 달성하여 최신 기술 수준을 확립한다.

ABSTRACT

Audio and visual signals complement each other in human speech perception, so do they in speech recognition. The visual hint is less evident than the acoustic hint, but more robust in a complex acoustic environment, as far as speech perception is concerned. It remains a challenge how we effectively exploit the interaction between audio and visual signals for automatic speech recognition. There have been studies to exploit visual signals as redundant or complementary information to audio input in a synchronous manner. Human studies suggest that visual signal primes the listener in advance as to when and on which frequency to attend to. We propose a Predict-and-Update Network (P&U net), to simulate such a visual cueing mechanism for Audio-Visual Speech Recognition (AVSR). In particular, we first predict the character posteriors of the spoken words, i.e. the visual embedding, based on the visual signals. The audio signal is then conditioned on the visual embedding via a novel cross-modal Conformer, that updates the character posteriors. We validate the effectiveness of the visual cueing mechanism through extensive experiments. The proposed P&U net outperforms the state-of-the-art AVSR methods on both LRS2-BBC and LRS3-BBC datasets, with the relative reduced Word Error Rate (WER)s exceeding 10% and 40% under clean and noisy conditions, respectively.

연구 동기 및 목표

  • 청각 품질이 악화되는 노이즈가 있는 환경에서의 강건한 음성 인식 도전 과제를 해결하기 위해.
  • 입술 움직임이 청각 처리를 예측하고 이끄는 인간 유사한 시각 단서 기반 메커니즘을 모델링하기 위해.
  • 단순한 후기 또는 초기 융합이 아닌, 예측적 시각 프라밍을 통합한 초기 융합 방식을 통해 음성-시각 음성 인식을 향상시키기 위해.
  • 시각 단서가 저 SNR 조건에서 특히 청각 신호가 손상된 경우에 강건성을 향상시킨다는 것을 검증하기 위해.
  • 음성-시각 상호작용이 단순한 신호 보완보다 더 효과적이라는 것을 입증하기 위해.

제안 방법

  • P&U net은 시각 인코더를 사용해 시각 신호에서 문자 사후확률을 예측하여, 예측적 단서로 기능하는 시각 임베딩을 생성한다.
  • 이 시각 임베딩은 새로운 크로스모달 Conformer 블록을 통해 청각 스트림을 조절하며, 인자화된 Excitation FFN을 통해 청각 표현을 업데이트한다.
  • 표준 FFN 레이어를 대체하여 동적이고 모odal 인식 기반의 특징 재조정이 가능한 인자화된 Excitation 메커니즘을 도입함으로써 크로스모달 Conformer를 설계한다.
  • 모델은 청각 처리 이전에 시각 예측을 통합함으로써 초기 융합을 구현하며, 인간 뇌의 예측적 주의 메커니즘을 시뮬레이션한다.
  • 사전 훈련된 단모달 모델을 초기화에 사용하고, 문자 수준의 사후확률에 대해 교차 엔트로피 손실을 사용하여 템플릿-포싱 방식으로 훈련한다.
  • 다양한 SNR 조건에서의 모달리티 정렬 분석을 위해 청각-시각 컨텍스트 표현 간 코사인 유사도를 사용한다.
Figure 1: The universal diagram of end-to-end lip reading or ASR . Input can be video or audio.
Figure 1: The universal diagram of end-to-end lip reading or ASR . Input can be video or audio.

실험 결과

연구 질문

  • RQ1시각 신호가 보완적 입력 외에도 음성-시각 음성 인식(AVSR)에서 청각 처리를 이끄는 예측적 단서로 사용될 수 있는가?
  • RQ2청각 입력을 예측하는 시각 단서 메커니즘이 노이즈가 많은 환경에서 인식의 강건성을 향상시키는가?
  • RQ3시각 조건부 처리가 네트워크 아키텍처 내에서 어떤 위치에 위치하는가에 따라 성능에 영향을 미치는가?
  • RQ4시각 임베딩의 차원이 인식 정확도에 어느 정도의 영향을 미치는가?
  • RQ5AVSR에서 표준 초기 또는 후기 융합 전략에 비해 시각 단서 기반 메커니즘이 더 효과적인가?

주요 결과

  • P&U net은 LRS2-BBC 데이터셋에서 청정 조건에서 최신 기술 수준의 방법에 비해 10퍼센트 이상의 상대적 WER 감소를 달성했고, 노이즈가 있는 조건에서는 40퍼센트 이상의 감소를 기록했다.
  • 저 SNR 조건에서 청각-시각 표현 간 코사인 유사도가 더 높게 유지됨으로써, P&U net은 모달리티 정렬이 더 우수하다는 것을 보여주었다.
  • P&U net (후기) 버전은 노이즈가 많은 조건에서 뛰어난 강건성을 보였으며, SNR이 감소함에 따라 코사인 유사도 감소 속도가 기준 모델보다 둔화되었다.
  • 시각 임베딩의 차원(K=16에서 128)이 성능에 미치는 영향은 미미하여, 중간 크기의 임베딩만으로도 효과적인 단서 기반 조건부 처리가 가능하다는 것을 시사한다.
  • 첫 번째 FFN 레이어에 인자화된 Excitation 모듈을 도입할 경우, 두 번째 또는 둘 다를 대체하는 것보다 더 높은 성능을 기록했으며, 이는 초기 단계에서의 시각 단서 기반 조건부 처리의 유용성을 확인한다.
  • 두 FFN 레이어 모두를 인자화된 Excitation 모듈로 대체할 경우 성능이 저하되었으며, 이는 시각 단서 기반 조건부 처리를 전역적으로 적용하기보다는 선택적으로 적용해야 한다는 것을 시사한다.
Figure 2: The general block diagram of our proposed P&U net which adopts video modality to predict a sequence of a coarse probability distribution, i.e., visual embedding, towards the target text transcription. The audio input is then augmented with the visual embedding for speech recognition via an
Figure 2: The general block diagram of our proposed P&U net which adopts video modality to predict a sequence of a coarse probability distribution, i.e., visual embedding, towards the target text transcription. The audio input is then augmented with the visual embedding for speech recognition via an

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.