[Paper Review] JVS corpus: free Japanese multi-speaker voice corpus
The paper introduces the JVS corpus, a free, 30-hour Japanese multi-speaker voice dataset in four sub-corpora (parallel100, nonpara30, whisper10, falsetto10) designed for multi-speaker and multi-style speech research.
Thanks to improvements in machine learning techniques, including deep learning, speech synthesis is becoming a machine learning task. To accelerate speech synthesis research, we are developing Japanese voice corpora reasonably accessible from not only academic institutions but also commercial companies. In 2017, we released the JSUT corpus, which contains 10 hours of reading-style speech uttered by a single speaker, for end-to-end text-to-speech synthesis. For more general use in speech synthesis research, e.g., voice conversion and multi-speaker modeling, in this paper, we construct the JVS corpus, which contains voice data of 100 speakers in three styles (normal, whisper, and falsetto). The corpus contains 30 hours of voice data including 22 hours of parallel normal voices. This paper describes how we designed the corpus and summarizes the specifications. The corpus is available at our project page.
Motivation & Objective
- Provide a large, high-quality, multi-speaker Japanese voice corpus for research and development in speech synthesis, voice conversion, and multi-speaker modeling.
- Offer parallel and non-parallel utterances to support diverse tasks such as voice conversion, speaker factorization, and multi-speaker modeling.
- Ensure accessibility and documentation through thorough annotations, licensing, and an easily downloadable format.
- Support both academic and commercial research use under clear licensing terms.
Proposed method
- Design four sub-corpora with specified utterance counts per speaker: parallel100 (100 parallel normal utterances), nonpara30 (30 non-parallel normal utterances), whisper10 (10 whisper utterances), and falsetto10 (10 falsetto utterances).
- Record 100 native Japanese professional speakers in a studio at 24 kHz, 16-bit RIFF WAV, with UTF-8 transcriptions and phoneme alignments.
- Automatically generate full context and monophone labels (Open JTalk) and align phonemes (Julius), with manually annotated F0 ranges per speaker.
- Provide additional tags: speaker similarity matrices, duration data, and per-speaker gender and F0 range information.
- Make the corpus freely available for research with clear licensing terms for academic and commercial use.
- Include transcriptions, phoneme alignments, and speaker-related metadata to facilitate tasks like voice conversion and multi-speaker modeling.
Experimental results
Research questions
- RQ1What is the structure and scope of the JVS corpus and how can it support multi-speaker and multi-style speech research?
- RQ2What are the specifications (speakers, utterances, styles, annotations) of the JVS corpus and how is data organized?
- RQ3How much data exists per sub-corpus and per speaker, and what are the recording/annotation pipelines?
- RQ4How can the JVS corpus enable research tasks such as voice conversion, speaker modeling, and style adaptation?
- RQ5What licensing and accessibility terms govern use of the JVS corpus for different research contexts?
Key findings
- The JVS corpus comprises 100 native Japanese professional speakers and about 30 hours of data across four sub-corpora.
- Parallel100 contains 100 utterances per speaker, while nonpara30, whisper10, and falsetto10 add non-parallel and style-varied data totaling 30.4 hours.
- Average normal-voice duration per speaker is about 15.7 minutes, with roughly 1.24 minutes of whisper and 1.18 minutes of falsetto per speaker.
- Speakers’ F0 ranges are manually annotated, and perceptual speaker similarity matrices are provided for cross-speaker analysis.
- Audio is 24 kHz, 16-bit RIFF WAV, with OpenJTalk and Julius-based labeling pipelines for full annotations.
- The corpus is free for academic and commercial research, with project-page licensing detailing commercial-use terms.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.