Skip to main content
QUICK REVIEW

[Paper Review] CoVoST 2: A Massively Multilingual Speech-to-Text Translation Corpus

Changhan Wang, Anne Wu|arXiv (Cornell University)|Jul 20, 2020
Natural Language Processing Techniques40 citations
TL;DR

CoVoST 2 is a large-scale, open-source multilingual speech-to-text translation corpus comprising 21 source languages to English and English to 15 target languages, making it the largest such dataset to date. It enables extensive training and evaluation of speech translation models across diverse language pairs, with data quality validated through sanity checks and released under CC0.

ABSTRACT

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to foster research in massive multilingual speech translation and speech translation for low resource language pairs, we release CoVoST 2, a large-scale multilingual speech translation corpus covering translations from 21 languages into English and from English into 15 languages. This represents the largest open dataset available to date from total volume and language coverage perspective. Data sanity checks provide evidence about the quality of the data, which is released under CC0 license. We also provide extensive speech recognition, bilingual and multilingual machine translation and speech translation baselines.

Motivation & Objective

  • To address the scarcity of large-scale, multilingual speech-to-text translation datasets, especially for low-resource language pairs.
  • To support research in massive multilingual speech translation by providing a dataset with broad language coverage and high data volume.
  • To improve model generalization and performance across diverse linguistic and phonetic variations through diverse multilingual data.
  • To enable benchmarking of speech recognition, machine translation, and end-to-end speech translation models across multiple language pairs.

Proposed method

  • The dataset is constructed from existing multilingual speech and text resources, including Common Voice and OPUS datasets, with automatic alignment of audio and text using forced alignment techniques.
  • Audio and text pairs are curated and filtered through automated data sanity checks to ensure quality and consistency.
  • The corpus includes parallel data for 21 source-to-English and 15 English-to-target language pairs, covering a wide range of linguistic diversity.
  • Baseline models are trained for speech recognition, bilingual and multilingual machine translation, and end-to-end speech translation using the dataset.
  • The dataset is released under the CC0 license to maximize accessibility and reuse in research.
  • Evaluation protocols and metrics are established to support consistent benchmarking across speech translation, translation, and ASR tasks.

Experimental results

Research questions

  • RQ1How does model performance in speech translation vary across high-resource and low-resource language pairs in a massively multilingual setting?
  • RQ2To what extent does multilingual pretraining on CoVoST 2 improve zero-shot transfer performance to unseen language pairs?
  • RQ3Can multilingual speech translation models trained on CoVoST 2 generalize effectively across diverse linguistic typologies and phonetic structures?
  • RQ4How does the inclusion of 21 source languages and 15 target languages impact the scalability and efficiency of end-to-end speech translation systems?
  • RQ5What are the key quality and coverage trade-offs in constructing a large-scale, open multilingual speech translation corpus?

Key findings

  • CoVoST 2 is the largest publicly available multilingual speech-to-text translation dataset, with comprehensive coverage of 21 source-to-English and 15 English-to-target language pairs.
  • Data sanity checks confirm high-quality alignment between audio and text, supporting reliable model training and evaluation.
  • The dataset enables training of strong multilingual speech translation baselines across diverse language pairs, including low-resource combinations.
  • Baseline models for speech recognition, machine translation, and speech translation show competitive performance, demonstrating the dataset's utility for benchmarking.
  • The release of the dataset under CC0 facilitates broad reuse and reproducibility in multilingual speech translation research.
  • The extensive language coverage supports the development of models with improved zero-shot and few-shot generalization across language pairs.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.