Skip to main content
QUICK REVIEW

[논문 리뷰] CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition

Linhao Dong, Bo Xu|arXiv (Cornell University)|2019. 05. 27.
Speech Recognition and Synthesis참고 문헌 40인용 수 6
한 줄 요약

이 논문은 종단 간 음성 인식을 위한 새로운 소프트이고 단조적인 어텐션 메커니즘인 연속적 통합-화염(CIF)을 제안한다. 이는 스파iking 신경망의 동역학을 모방하며, 음성 특징을 연속적으로 통합하고 감지된 경계에서 화염을 발생시음으로써 효율적인 온라인 추론과 정확한 음향 경계 국소화를 가능하게 한다. 이로 인해 Librispeech test-clean에서 새로운 최고 성능인 2.86%의 WER과 HKUST 매파오어 전화 음성 인식에서 23.09%의 CER를 달성하였다.

ABSTRACT

In this paper, we propose a novel soft and monotonic alignment mechanism used for sequence transduction. It is inspired by the integrate-and-fire model in spiking neural networks and employed in the encoder-decoder framework consists of continuous functions, thus being named as: Continuous Integrate-and-Fire (CIF). Applied to the ASR task, CIF not only shows a concise calculation, but also supports online recognition and acoustic boundary positioning, thus suitable for various ASR scenarios. Several support strategies are also proposed to alleviate the unique problems of CIF-based model. With the joint action of these methods, the CIF-based model shows competitive performance. Notably, it achieves a word error rate (WER) of 2.86% on the test-clean of Librispeech and creates new state-of-the-art result on Mandarin telephone ASR benchmark.

연구 동기 및 목표

  • 표준 어텐션 메커니즘이 종단 간 음성 인식에서 겪는 한계, 즉 온라인 추론 지원 부족과 프레임 동기화 정렬의 부재를 해결하기 위해.
  • 실시간 스트리밍과 음향 경계 탐지를 지원하는 소프트이고 단조적인 정렬 메커니즘을 개발하기 위해.
  • 시퀀스 변환을 위한 이산적 통합-화염 스파iking 뉴런의 대체로, 미분 가능한 연속 시간 기반의 방법을 설계하기 위해.
  • 스케일링, 양상 손실, 꼬리 처리와 같은 새로운 보조 전략을 통해 모델의 효율성과 정렬 정확도를 향상시키기 위해.

제안 방법

  • CIF는 시간에 따라 인코더 표현의 연속적 통합을 모델링하며, 막 전위는 입력 가중치와 은닉 상태에 기반해 누적된다.
  • 통합된 값이 임계값에 도달하면 '스パイ크'가 생성되며, 이는 현재 레이블에 대한 음향 정보를 출력하고 다음 레이블로 분할하는 데를 유도한다.
  • 이 메커니즘은 연속 함수를 사용하여 통합-화염 과정을 시뮬레이션하여 시간에 따른 역전파를 가능하게 한다.
  • 스케일링 전략은 교차 엔트로피 학습 중 레이블-대상 길이 불일치를 조정하기 위해 사용된다.
  • 양상 손실은 예측된 출력 토큰 수를 감독하기 위해 도입된다.
  • 꼬리 처리 방법은 시퀀스의 끝부분에 잔류하는 정보를 처리하여 정렬 정확도를 향상시킨다.
Fig. 1 : Illustration of the attention alignment and our proposed CIF alignment on an encoded utterance of length 5 and labelled as ”CAT”. The shade of gray in each square represents the weight of each encoder step involved in the calculation of decoding labels. The vertically dashed line in (b) rep
Fig. 1 : Illustration of the attention alignment and our proposed CIF alignment on an encoded utterance of length 5 and labelled as ”CAT”. The shade of gray in each square represents the weight of each encoder step involved in the calculation of decoding labels. The vertically dashed line in (b) rep

실험 결과

연구 질문

  • RQ1온라인 음성 인식을 지원하면서도 높은 정확도를 유지할 수 있는, 미분 가능하고 연속 시간 기반의 어텐션 메커니즘을 설계할 수 있는가?
  • RQ2스파iking 신경망의 통합-화염 원리를 종단 간 ASR에 적응시켜 역전파가 가능한 방식으로 어떻게 활용할 수 있는가?
  • RQ3CIF 기반 모델의 안정성과 성능 향상을 위해 필요한 보조 학습 전략은 무엇인가?
  • RQ4CIF는 자동 음성 인식과 음향 경계 탐지 모두에서 기존의 단조적이고 소프트 어텐션 메커니즘을 초월하는가?
  • RQ5CIF는 독서 및 즉흥적 말하기와 같은 다양한 음성 인식 작업에 일반화 가능한가?

주요 결과

  • CIF 기반 모델은 Librispeech test-clean 세트에서 기존 방법을 뛰어넘는 새로운 최고 성능인 2.86%의 WER를 달성하였다.
  • AISHELL-2 매파오어 독서 음성 벤치마크에서 CIF는 test_android에서 6.17%의 CER를 기록하여, Chain-TDNN 기반 모델의 9.59%보다 뚜렷한 향상을 이뤘다.
  • 더 도전적인 HKUST 매파오어 전화 대화 데이터셋에서는 CIF가 23.09%의 CER를 기록하며 새로운 최고 성능을 수립했으며, Transformer와 Joint CTC-attention 모델을 모두 앞섰다.
  • 절단 실험 결과, 양상 손실이 가장 중요한 구성 요소로 밝혀졌으며, 이 손실을 제거할 경우 성능 저하와 학습 불안정성이 가장 크게 발생했다.
  • 청각 기반 추론을 청크 히프팅 방식으로 지원하여, 스트리밍 기능을 갖춘 상태에서 test-clean에서 3.96%의 WER를 달성하였으며, 실용적 적용 가능성을 입증하였다.
  • 자동 회귀 디코더는 경계 명료도가 낮은 상황에서 가장 중요한 역할을 하지만, 매파오어와 같은 고명료도 설정에서도 효과적으로 기능한다.
Fig. 2 : The architecture of our CIF-based model used for the ASR task. Operations in the dashed rectangles are only applied in the training stage. The switch (S) before the CIF module connects the left in the training stage and the right in the inference stage.
Fig. 2 : The architecture of our CIF-based model used for the ASR task. Operations in the dashed rectangles are only applied in the training stage. The switch (S) before the CIF module connects the left in the training stage and the right in the inference stage.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.