[논문 리뷰] Jira: a Kurdish Speech Recognition System Designing and Building Speech Corpus and Pronunciation Lexicon
이 논문은 제어된 환경과 커뮤니티 기반 환경에서 576명의 참가자로부터 43.68시간의 음성 데이터를 수집하고, 60만 단어의 발음 어휘사전을 구축하여, 중앙 Kurdisht를 위한 첫 번째 대규모 어휘 어휘 인식 시스템인 Jira를 제안한다. Kaldi 툴킷을 사용하여, 다양한 주제에서 13.9%의 단어 오류율을 달성했고, 일반 주제에서는 4.9%로 낮아져 Kurdisht NLP 자원의 기초적 기여를 이룬다.
In this paper, we introduce the first large vocabulary speech recognition system (LVSR) for the Central Kurdish language, named Jira. The Kurdish language is an Indo-European language spoken by more than 30 million people in several countries, but due to the lack of speech and text resources, there is no speech recognition system for this language. To fill this gap, we introduce the first speech corpus and pronunciation lexicon for the Kurdish language. Regarding speech corpus, we designed a sentence collection in which the ratio of di-phones in the collection resembles the real data of the Central Kurdish language. The designed sentences are uttered by 576 speakers in a controlled environment with noise-free microphones (called AsoSoft Speech-Office) and in Telegram social network environment using mobile phones (denoted as AsoSoft Speech-Crowdsourcing), resulted in 43.68 hours of speech. Besides, a test set including 11 different document topics is designed and recorded in two corresponding speech conditions (i.e., Office and Crowdsourcing). Furthermore, a 60K pronunciation lexicon is prepared in this research in which we faced several challenges and proposed solutions for them. The Kurdish language has several dialects and sub-dialects that results in many lexical variations. Our methods for script standardization of lexical variations and automatic pronunciation of the lexicon tokens are presented in detail. To setup the recognition engine, we used the Kaldi toolkit. A statistical tri-gram language model that is extracted from the AsoSoft text corpus is used in the system. Several standard recipes including HMM-based models (i.e., mono, tri1, tr2, tri2, tri3), SGMM, and DNN methods are used to generate the acoustic model. These methods are trained with AsoSoft Speech-Office and AsoSoft Speech-Crowdsourcing and a combination of them. The best performance achieved by the SGMM acoustic model which results in 13.9% of the average word error rate (on different document topics) and 4.9% for the general topic.
연구 동기 및 목표
- 중앙 Kurdisht에 대한 음성 및 텍스트 자원의 부족을 해결하기 위해, 약 3,000만 명이 사용하는 언어이지만 자원이 부족한 상황을 개선한다.
- 실제 중앙 Kurdisht 음성의 특성을 반영한 이음소음 분포를 갖춘 음성 데이터 코퍼스를 설계한다.
- 다양한 Kurdisht 방언 간의 어휘 변형을 고려한 표준화된 대규모 발음 어휘사전을 개발한다.
- 최신 음성 인식 기술을 활용하여 중앙 Kurdisht를 위한 대규모 어휘 음성 인식 시스템(LVSR)을 구축하고 평가한다.
- 재사용 가능한 음성 및 언어 자원을 제공함으로써 Kurdisht의 기초 NLP 인프라를 구축한다.
제안 방법
- 실제 중앙 Kurdisht 데이터의 이음소음 빈도를 반영하도록 문장 수집을 설계하여 언어학적 대표성을 확보한다.
- AsoSoft Speech-Office(제어된, 소음이 없는 환경)와 AsoSoft Speech-Crowdsourcing(텔레그램을 통한 모바일 기기 활용)를 통해 총 576명의 참가자로부터 43.68시간의 음성 데이터를 수집한다.
- 평가의 강건성을 확보하기 위해, 11개의 문서 주제를 포함한 테스트 세트를 구축하였으며, 사무실 및 커뮤니티 기반 환경에서 모두 녹음하였다.
- Kurdisht 방언 간의 철자 변형을 표준화하는 기법을 사용하여 60만 단어의 발음 어휘사전을 구축하였다.
- 어휘 사전 항목의 청각적 발음 표기를 생성하기 위해 자동 발음 모델링 기법을 적용하였다.
- Kaldi 툴킷을 사용하여, 두 가지 음성 수집 환경의 데이터를 통합하여, HMM 기반 모델(Mono, Tri1–Tri3), SGMM, DNN 등 다양한 음성 모델을 훈련시켰다.
실험 결과
연구 질문
- RQ1환경적 소음이 최소화되고 실제 세계의 변동성을 반영하는 대규모 언어학적으로 대표적인 중앙 Kurdisht 음성 코퍼스를 어떻게 수집할 수 있는가?
- RQ2Kurdisht 방언 간의 철자 변형을 효과적으로 표준화하고, 대규모 어휘사전을 위한 정확한 발음을 어떻게 생성할 수 있는가?
- RQ3제어된 환경과 커뮤니티 기반 음성 데이터를 결합할 경우, Kurdisht LVSR 시스템의 성능에 어떤 영향을 미치는가?
- RQ4저자원 환경에서 낮은 단어 오류율을 달성하기 위해 최적의 음성 모델링 기법(예: SGMM, DNN)은 무엇인가?
- RQ5AsoSoft 텍스트 코퍼스 기반의 삼중어휘 언어 모델은 다양한 주제에서 인식 성능을 얼마나 향상시키는가?
주요 결과
- AsoSoft Speech-Office와 AsoSoft Speech-Crowdsourcing 데이터셋은 총 43.68시간의 고품질이며 다양한 음성 데이터를 576명의 참가자로부터 확보하였다.
- SGMM 음성 모델을 사용하여 11개의 다른 문서 주제 평균에서 13.9%의 단어 오류율을 달성하였다.
- 일반 주제에서는 단어 오류율이 4.9%로 감소하여, 광범위하고 도메인 특화가 아닌 음성에서 뛰어난 성능을 보였다.
- 철자 표준화 및 자동 발음 모델링 기법을 활용하여 60만 단어의 발음 어휘사전을 성공적으로 구축하였으며, 방언 간의 변형 문제를 해결하였다.
- SGMM 모델은 HMM 및 DNN 베이스라인을 모두 능가하여, 저자원 Kurdisht 음성 인식에 효과적임을 입증하였다.
- AsoSoft 텍스트 코퍼스 기반의 삼중어휘 언어 모델 통합으로 인해, 다양한 주제에서의 인식 정확도가 크게 향상되었다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.