Skip to main content
QUICK REVIEW

[논문 리뷰] Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents

Priyank Dubey, Bilal Shah|arXiv (Cornell University)|2022. 04. 03.
Speech Recognition and Synthesis인용 수 14
한 줄 요약

이 논문은 인도-영어 발음에 적합한 테스트된 DeepSpeech 기반 엔드 투 엔드 ASR 시스템을 제안한다. 전이 학습과 Indic TTS 데이터셋을 이용한 데이터 증강 기법을 활용하여, 사전 훈련된 DeepSpeech-0.9.3 모델을 인도-영어 음성에 맞게 보정한다. 사전 훈련된 모델을 인도-영어 음성에 맞게 보정함으로써, 미훈련 모델 및 상용 서비스보다 뚜렷한 성능 향상을 이룩하였으며, 인도 내 다양한 지역적 영어 발음에 대해 강력한 성능을 보여주었다.

ABSTRACT

Automated Speech Recognition (ASR) is an interdisciplinary application of computer science and linguistics that enable us to derive the transcription from the uttered speech waveform. It finds several applications in Military like High-performance fighter aircraft, helicopters, air-traffic controller. Other than military speech recognition is used in healthcare, persons with disabilities and many more. ASR has been an active research area. Several models and algorithms for speech to text (STT) have been proposed. One of the most recent is Mozilla Deep Speech, it is based on the Deep Speech research paper by Baidu. Deep Speech is a state-of-art speech recognition system is developed using end-to-end deep learning, it is trained using well-optimized Recurrent Neural Network (RNN) training system utilizing multiple Graphical Processing Units (GPUs). This training is mostly done using American-English accent datasets, which results in poor generalizability to other English accents. India is a land of vast diversity. This can even be seen in the speech, there are several English accents which vary from state to state. In this work, we have used transfer learning approach using most recent Deep Speech model i.e., deepspeech-0.9.3 to develop an end-to-end speech recognition system for Indian-English accents. This work utilizes fine-tuning and data argumentation to further optimize and improve the Deep Speech ASR system. Indic TTS data of Indian-English accents is used for transfer learning and fine-tuning the pre-trained Deep Speech model. A general comparison is made among the untrained model, our trained model and other available speech recognition services for Indian-English Accents.

연구 동기 및 목표

  • 기존 ASR 시스템이 주로 미국-영어 데이터로 훈련되어 인도-영어 발음에 대해 일반화 능력이 떨어지는 문제를 해결하기 위해.
  • 인도-영어 발음 패턴의 언어적 다양성에 맞게 최적화된 강력한 엔드 투 엔드 ASR 시스템을 개발하기 위해.
  • 사전 훈련된 DeepSpeech-0.9.3 모델을 활용한 전이 학습을 통해 훈련 데이터 및 계산 자원 요구량을 줄이기 위해.
  • Indic TTS 데이터셋을 기반으로 한 세부 조정 및 데이터 증강을 통해 모델 성능을 최적화하기 위해.

제안 방법

  • 인도-영어 발음을 포함한 Indic TTS 데이터셋을 사용하여 사전 훈련된 DeepSpeech-0.9.3 모델을 전이 학습을 통해 보정함.
  • 훈련 데이터의 다양성을 높이고 모델의 강건성을 향상시키기 위해 데이터 증강 기법을 적용함.
  • 연결성 시간 분류(CTC) 손실을 사용하는 깊이 있는 양방향 장기 단기 기억(비-LSTM) 네트워크 기반의 순서-순서 아키텍처를 활용함.
  • 수렴 속도를 높이고 최적화 효율성을 향상시키기 위해 다중 GPU를 사용하여 모델을 훈련함.
  • 미훈련 모델, 제안된 모델 및 상용 ASR 서비스 간의 단어 오류율(WER) 비교를 통해 성능을 평가함.

실험 결과

연구 질문

  • RQ1제한된 도메인 전용 데이터로 인도-영어 발음에 효과적으로 적응할 수 있는 사전 훈련된 DeepSpeech 모델을 전이 학습으로 활용할 수 있는가?
  • RQ2데이터 증강은 인도-영어 음성에 대한 ASR 모델의 강건성과 정확도에 어떤 영향을 미치는가?
  • RQ3세부 조정된 DeepSpeech 모델은 인도-영어 음성 인식 작업에서 미훈련 모델 대비 어떤 성능 향상을 보이는가?
  • RQ4제안된 시스템은 기존 상용 ASR 서비스와 비교해 인도-영어 발음에 대해 얼마나 정확한가?

주요 결과

  • 세부 조정된 DeepSpeech 모델은 인도-영어 음성에서 미훈련 기준 모델 대비 단어 오류율(WER)이 뚜렷이 감소함.
  • 제안된 시스템은 인도-영어 발음에 대해 상용 ASR 서비스를 뛰어넘는 WER 성능을 보이며, 도메인 특화 적응 능력이 뛰어남을 시사함.
  • 데이터 증강은 특히 인도 내 드문 발음이나 자원이 적은 지역 발음에 대해 모델의 일반화 능력을 향상시키는 데 기여함.
  • 전이 학습을 통해 사전 훈련된 모델을 최소한의 추가 훈련 데이터와 계산 비용으로 효과적으로 적응시킬 수 있음.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.