Skip to main content
QUICK REVIEW

[Paper Review] KazakhTTS2: Extending the Open-Source Kazakh TTS Corpus With More Data, Speakers, and Topics

Saida Mussakhojayeva, Yerbolat Khassanov|arXiv (Cornell University)|Jan 15, 2022
Speech Recognition and Synthesis4 citations
TL;DR

This paper presents KazakhTTS2, an expanded open-source text-to-speech (TTS) corpus for Kazakh, increasing from 93 to 271 hours of transcribed audio across five speakers (three female, two male) and diverse sources including news, books, and Wikipedia. The corpus enables training of high-quality neural TTS models, achieving subjective mean opinion scores (MOS) between 3.6 and 4.2, demonstrating its suitability for real-world applications and advancing low-resource Turkic language research.

ABSTRACT

We present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two to five (three females and two males), and the topic coverage has been diversified with the help of new sources, including a book and Wikipedia articles. This corpus is necessary for building high-quality TTS systems for Kazakh, a Central Asian agglutinative language from the Turkic family, which presents several linguistic challenges. We describe the corpus construction process and provide the details of the training and evaluation procedures for the TTS system. Our experimental results indicate that the constructed corpus is sufficient to build robust TTS models for real-world applications, with a subjective mean opinion score ranging from 3.6 to 4.2 for all the five speakers. We believe that our corpus will facilitate speech and language research for Kazakh and other Turkic languages, which are widely considered to be low-resource due to the limited availability of free linguistic data. The constructed corpus, code, and pretrained models are publicly available in our GitHub repository.

Motivation & Objective

  • To address the lack of large-scale, high-quality, open-source speech corpora for Kazakh, a low-resource, agglutinative Turkic language.
  • To enhance the original KazakhTTS corpus by increasing data volume, speaker diversity, and topic coverage.
  • To validate the corpus’s utility for training robust, deployable TTS systems through subjective evaluation.
  • To support future research in Kazakh and other low-resource Turkic languages via open access to data, code, and pretrained models.

Proposed method

  • Collected and manually transcribed 271 hours of audio from diverse sources: news articles, a published book, and Wikipedia content.
  • Selected five professional speakers (three female, two male) to ensure high audio quality and linguistic diversity.
  • Constructed a neural TTS system using the Tacotron 2 architecture for model training and evaluation.
  • Conducted crowdsourced subjective evaluation using Mean Opinion Score (MOS) to assess naturalness and quality of synthesized speech.
  • Performed detailed error analysis on synthesized outputs to identify common failure modes such as mispronunciation and word skipping.
  • Released the full corpus, training code, and pretrained models via a public GitHub repository for academic and commercial use.

Experimental results

Research questions

  • RQ1Can a significantly expanded and diversified TTS corpus improve the quality and robustness of Kazakh text-to-speech synthesis?
  • RQ2To what extent does increasing speaker diversity and data volume enhance the naturalness and intelligibility of synthesized speech in Kazakh?
  • RQ3How do different text sources (news, books, Wikipedia) affect the performance and generalization of TTS models?
  • RQ4What are the primary error types in synthesized speech, and how can they be mitigated in future models?
  • RQ5Can the KazakhTTS2 corpus serve as a foundation for transfer learning and cross-lingual applications in other Turkic languages?

Key findings

  • The KazakhTTS2 corpus comprises 271 hours of high-quality, manually transcribed audio from five professional speakers across diverse sources.
  • All five speakers achieved subjective Mean Opinion Scores (MOS) between 3.6 and 4.2, indicating high speech quality suitable for real-world deployment.
  • Speaker M2 achieved the highest MOS (4.2), while F1 and M1 News had the highest scores in the prior release (4.726 and 4.360, respectively).
  • The most frequent error types in synthesized speech were mispronunciation (15 errors), incomplete words (14), and skipped words (14), with M1 showing the highest error count.
  • The corpus enables training of state-of-the-art TTS models using Tacotron 2, with results consistent across multiple evaluation metrics and confidence intervals.
  • The corpus, code, and pretrained models are publicly available on GitHub, supporting reproducibility and future research in Kazakh and related low-resource languages.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.