Skip to main content
QUICK REVIEW

[Paper Review] JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

Ryosuke Sonobe, Shinnosuke Takamichi|arXiv (Cornell University)|Oct 28, 2017
Speech Recognition and Synthesis6 references88 citations
TL;DR

The paper introduces the JSUT corpus, a free 10-hour Japanese speech dataset designed to cover all daily-use kanji pronunciations for end-to-end speech synthesis, with nine sub-corpora and publicly available resources.

ABSTRACT

Thanks to improvements in machine learning techniques including deep learning, a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role. However, such a corpus for Japanese speech synthesis does not exist. In this paper, we designed a novel Japanese speech corpus, named the "JSUT corpus," that is aimed at achieving end-to-end speech synthesis. The corpus consists of 10 hours of reading-style speech data and its transcription and covers all of the main pronunciations of daily-use Japanese characters. In this paper, we describe how we designed and analyzed the corpus. The corpus is freely available online.

Motivation & Objective

  • Provide a free, large-scale Japanese speech corpus for end-to-end speech synthesis research.
  • Ensure coverage of all pronunciations of daily-use Japanese characters and individual readings.
  • Include diverse domains such as loanwords, paraphrase variants, and travel/precedent content.

Proposed method

  • Design nine sub-corpora to cover main pronunciations and readings of daily-use kanji characters.
  • Construct basic5000 by selecting 5000 sentences from Wikipedia and TANAKA corpus plus manually added sentences.
  • Crowdsource or collect utterances including counter suffixes, loanwords, paraphrases, and domain-specific content.
  • Incorporate para-speech from voice actress corpus and onomatopoeia sentences to enrich prosody.
  • Record 10 hours of read speech from a native female speaker in an anechoic room at 48 kHz.
  • Annotate and analyze linguistic statistics such as moras and words per utterance and measure F0 variation across days.

Experimental results

Research questions

  • RQ1Can a freely accessible Japanese speech corpus be constructed to cover all pronunciations of daily-use kanji characters for end-to-end TTS?
  • RQ2How do domain variations (loanwords, paraphrase, travel, precedent, etc.) affect data diversity and pronunciation coverage?
  • RQ3What are the linguistic and speech statistics of the constructed JSUT corpus and how do they vary over recording days?

Key findings

  • The JSUT corpus comprises nine sub-corpora with varied linguistic content and 10 hours of speech data.
  • The basic5000 sub-corpus targets all main pronunciations of daily-use kanji characters.
  • The corpus includes loanwords, paraphrase variants, onomatopoeia, and domain-specific sentences to broaden coverage.
  • Speech data were recorded in an anechoic room at 48 kHz from a native female speaker.
  • The corpus provides UTF-8 text and 16-bit WAV speech data and is freely available online.
  • Analysis shows a range of utterance lengths and an observable increase in mean log F0 over recording days.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.