Skip to main content
QUICK REVIEW

[논문 리뷰] FluentNet: End-to-End Detection of Speech Disfluency with Deep Learning

Tedd Kourkounakis, Amirhossein Hajavi|arXiv (Cornell University)|2020. 09. 23.
Stuttering Research and Treatment참고 문헌 46인용 수 9
한 줄 요약

FluentNet는 스펙트럼 특징 학습을 위해 Squeeze-and-Excitation 잔차 CNN를, 시간적 모델링을 위해 양방향 LSTM을, 전역 주의 메커니즘을 사용하여 여섯 가지 유형의 언어 불순성—소리, 단어, 구절 반복, 수정, 중간 말하기, 연장—을 탐지하는 엔드 투 엔드 딥 러닝 모델이다. 이 모델은 UCLASS 데이터셋에서 최고 성능을 기록하며, 새로운 합성 데이터셋인 LibriStutter에도 잘 일반화된다.

ABSTRACT

Strong presentation skills are valuable and sought-after in workplace and classroom environments alike. Of the possible improvements to vocal presentations, disfluencies and stutters in particular remain one of the most common and prominent factors of someone's demonstration. Millions of people are affected by stuttering and other speech disfluencies, with the majority of the world having experienced mild stutters while communicating under stressful conditions. While there has been much research in the field of automatic speech recognition and language models, there lacks the sufficient body of work when it comes to disfluency detection and recognition. To this end, we propose an end-to-end deep neural network, FluentNet, capable of detecting a number of different disfluency types. FluentNet consists of a Squeeze-and-Excitation Residual convolutional neural network which facilitate the learning of strong spectral frame-level representations, followed by a set of bidirectional long short-term memory layers that aid in learning effective temporal relationships. Lastly, FluentNet uses an attention mechanism to focus on the important parts of speech to obtain a better performance. We perform a number of different experiments, comparisons, and ablation studies to evaluate our model. Our model achieves state-of-the-art results by outperforming other solutions in the field on the publicly available UCLASS dataset. Additionally, we present LibriStutter: a disfluency dataset based on the public LibriSpeech dataset with synthesized stutters. We also evaluate FluentNet on this dataset, showing the strong performance of our model versus a number of benchmark techniques.

연구 동기 및 목표

  • 다양한 유형의 언어 불순성을 정확하게 탐지하기 위한 엔드 투 엔드 딥 러닝 모델 개발.
  • LibriSpeech 기반으로 합성된 데이터셋인 LibriStutter를 통해 레이블이 부족한 언어 불순성 데이터 문제를 해결.
  • 단순한 채움말 탐지 초과 복잡한 스터티 패턴을 모델링하여 언어 불순성 탐지 성능 향상.
  • Squeeze-and-Excitation 및 주의 메커니즘과 같은 핵심 구성 요소가 언어 불순성 인식에 기여하는 방식 평가.
  • UCLASS와 신규 도입된 LibriStutter와 같은 공개 벤치마크에서 최고 수준의 성능 입증.

제안 방법

  • FluentNet는 원시 오디오에서 강건한 스펙트럼 프레임 수준 표현을 학습하기 위해 Squeeze-and-Excitation(SE) 잔차 컨volution 신경망을 사용한다.
  • 스터티된 언어에서 장거리 시간적 의존성을 모델링하기 위해 양방향 장기 단기 기억(LSTM, BLSTM) 레이어를 활용한다.
  • BLSTM 레이어 이후 전역 주의 메커니즘이 적용되어 중요도가 높은 언어 세그먼트에 집중함으로써 언어 불순성 분류 성능을 향상시킨다.
  • 학습은 교차 엔트로피 손실을 사용하고 학습률 $10^{-4}$로 Adam 최적화 기법을 통해 엔드 투 엔드로 수행된다.
  • UCLASS 데이터셋에서의 초모수 탐색을 통해 8개의 잔차 블록과 2개의 BLSTM 레이어로 아키텍처가 최적화되었다.
  • 합성된 언어 불순성 데이터셋인 LibriStutter는 깨끗한 LibriSpeech 오디오 샘플에 합성된 스터티를 주입하여 생성되었다.

실험 결과

연구 질문

  • RQ1엔드 투 엔드 딥 러닝 모델이 다양한 유형의 언어 불순성을 탐지하는 데 최고 성능을 낼 수 있는가?
  • RQ2Squeeze-and-Excitation 및 주의 메커니즘과 같은 구성 요소가 언어 불순성 탐지 성능에 기여하는 정도는 어떠한가?
  • RQ3실제 스터티 패턴을 모방하는 합성 언어 불순성 데이터에 대해 모델의 일반화 능력은 어느 정도인가?
  • RQ4수정 및 연장과 같은 드문 또는 복잡한 패턴을 포함한 다양한 언어 불순성 유형에 대해 모델의 성능는 어떠한가?
  • RQ5실제 데이터가 부족할 경우, LibriStutter와 같은 합성 데이터셋이 언어 불순성 탐지 모델 훈련에 실질적인 대안이 될 수 있는가?

주요 결과

  • FluentNet는 UCLASS 데이터셋에서 기존 방법들을 능가하는 최고 성능을 기록하며 언어 불순성 분류 성능을 확보했다.
  • UCLASS 데이터셋에서 모든 언어 불순성 유형에 대해 약 20 에포크 후 훈련 정확도가 거의 완벽해졌다.
  • 추론 실험 결과, Squeeze-and-Excitation 구성 요소를 제거할 경우 모든 언어 불순성 유형에서 정확도가 크게 떨어지고 빠짐 비율이 증가하는 것으로 나타났다.
  • 주의 메커니즘이 성능 향상에 크게 기여했으며, 이를 제거하면 분류 정확도에 명확한 하락이 관찰되었다.
  • 모델는 합성된 LibriStutter 데이터셋으로도 효과적으로 일반화되었으며, 성능 저하가 관찰되었지만 여전히 기준 성능를 뛰어넘었다.
  • 추론 결과는 UCLASS와 LibriStutter 양쪽에서 일관되게 나타났으며, 이는 합성 데이터셋의 타당성과 현실성에 대한 확인을 뒷받침한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.