Skip to main content
QUICK REVIEW

[Paper Review] A large-scale and PCR-referenced vocal audio dataset for COVID-19

Jobie Budd, Kieran Baker|arXiv (Cornell University)|Dec 15, 2022
Phonocardiography and Auscultation Techniques4 citations
TL;DR

This paper introduces the UK COVID-19 Vocal Audio Dataset, a large-scale, PCR-confirmed collection of vocal recordings (coughs, exhalations, speech) from 70,794 participants across England, linked to SARS-CoV-2 test results and self-reported symptoms. It enables machine learning models to detect infection status and respiratory conditions using audio, with 24,155 positive cases and 27.2% influenza test data, representing the largest such dataset to date.

ABSTRACT

The UK COVID-19 Vocal Audio Dataset is designed for the training and evaluation of machine learning models that classify SARS-CoV-2 infection status or associated respiratory symptoms using vocal audio. The UK Health Security Agency recruited voluntary participants through the national Test and Trace programme and the REACT-1 survey in England from March 2021 to March 2022, during dominant transmission of the Alpha and Delta SARS-CoV-2 variants and some Omicron variant sublineages. Audio recordings of volitional coughs, exhalations, and speech were collected in the 'Speak up to help beat coronavirus' digital survey alongside demographic, self-reported symptom and respiratory condition data, and linked to SARS-CoV-2 test results. The UK COVID-19 Vocal Audio Dataset represents the largest collection of SARS-CoV-2 PCR-referenced audio recordings to date. PCR results were linked to 70,794 of 72,999 participants and 24,155 of 25,776 positive cases. Respiratory symptoms were reported by 45.62% of participants. This dataset has additional potential uses for bioacoustics research, with 11.30% participants reporting asthma, and 27.20% with linked influenza PCR test results.

Motivation & Objective

  • To create a large-scale, clinically referenced vocal audio dataset to support machine learning models in detecting SARS-CoV-2 infection from voice samples.
  • To link vocal recordings with gold-standard SARS-CoV-2 PCR test results for reliable model training and evaluation.
  • To include self-reported symptoms, respiratory conditions, and influenza test results to enable broader bioacoustics and multi-condition research.
  • To support the development of non-invasive, scalable screening tools for respiratory infections using voice-based diagnostics.
  • To provide a publicly available, ethically sourced dataset representative of real-world population diversity during dominant SARS-CoV-2 variant circulation.

Proposed method

  • Participants were recruited via the UK Test and Trace programme and the REACT-1 survey between March 2021 and March 2022.
  • Vocal samples were collected through a digital survey, including volitional coughs, exhalations, and sustained speech.
  • Each participant’s audio was linked to their SARS-CoV-2 PCR test result, self-reported symptoms, and pre-existing respiratory conditions.
  • The dataset includes 70,794 participants with PCR-confirmed results and 24,155 positive cases, with 27.2% having concurrent influenza PCR results.
  • Demographic, symptom, and health history data were collected alongside audio to support multi-modal analysis.
  • The dataset was curated and released with ethical oversight, ensuring participant privacy and data integrity.

Experimental results

Research questions

  • RQ1Can vocal audio features alone reliably distinguish SARS-CoV-2 positive individuals from negative ones using machine learning?
  • RQ2To what extent do vocal patterns correlate with self-reported respiratory symptoms or pre-existing conditions like asthma?
  • RQ3How does the inclusion of influenza test results enhance the utility of the dataset for differential diagnosis research?
  • RQ4Can the dataset support the development of scalable, non-invasive screening tools for respiratory infections in community settings?
  • RQ5What is the impact of variant dynamics (Alpha, Delta, Omicron) on vocal biomarkers detectable in audio?

Key findings

  • The dataset comprises 70,794 participants with PCR-confirmed SARS-CoV-2 status, representing the largest such collection to date.
  • Of these, 24,155 tested positive for SARS-CoV-2, providing substantial data for training and evaluating diagnostic models.
  • Respiratory symptoms were reported by 45.62% of participants, indicating a strong link between vocal features and symptomatology.
  • 11.30% of participants reported asthma, enabling research into vocal biomarkers for chronic respiratory conditions.
  • 27.20% of participants had linked influenza PCR test results, supporting multi-pathogen diagnostic research.
  • The dataset spans the Alpha, Delta, and early Omicron variant periods, capturing vocal changes across dominant SARS-CoV-2 lineages.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.