[Paper Review] Audio, Speech, Language, & Signal Processing for COVID-19: A Comprehensive Overview.
This paper presents a comprehensive review of audio, speech, and signal processing techniques for COVID-19 screening, diagnosis, monitoring, and awareness. It synthesizes research on detecting COVID-19 markers in coughs, breathing patterns, and speech, leveraging AI to enable non-invasive, scalable digital health tools with demonstrated accuracy in symptom detection and emotional state assessment.
The Coronavirus (COVID-19) pandemic has been the research focus world-wide in the year 2020. Several efforts, from collection of COVID-19 patients' data to screening them for the virus's detection are taken with rigour. A major portion of COVID-19 symptoms are related to the functioning of the respiratory system, which in-turn critically influences the human speech production system. This drives the research focus towards identifying the markers of COVID-19 in speech and other human generated audio signals. In this paper, we give an overview of the speech and other audio signal, language and general signal processing-based work done using Artificial Intelligence techniques to screen, diagnose, monitor, and spread the awareness aboutCOVID-19. We also briefly describe the research related to detect accord-ing COVID-19 symptoms carried out so far. We aspire that this collective information will be useful in developing automated systems, which can help in the context of COVID-19 using non-obtrusive and easy to use modalities such as audio, speech, and language.
Motivation & Objective
- To consolidate and analyze existing research on using audio and speech signals for detecting and monitoring COVID-19 symptoms.
- To identify key AI-based signal processing techniques applicable to non-invasive, remote health monitoring during the pandemic.
- To evaluate the potential of speech and audio analysis in detecting respiratory and psychological markers of COVID-19.
- To highlight gaps in current research and guide future development of reliable, scalable digital health tools.
- To support the creation of automated, low-cost, and accessible systems for early detection and mental health support during pandemics.
Proposed method
- Systematic review of peer-reviewed literature and technical reports on audio, speech, and signal processing for COVID-19 applications.
- Categorization of studies based on application domains: screening/diagnosis, monitoring, awareness, and behavioral/mental health tracking.
- Analysis of machine learning and deep learning models used, including CNNs, RNNs, attention mechanisms, and feature extraction techniques like MFCCs and BoAW.
- Evaluation of model performance using standard metrics such as accuracy, F1-score, Spearman’s correlation, RMSE, and MAE across diverse datasets.
- Inclusion of benchmark datasets such as DAIC-WOZ, COVIG-19 Plasma Alliance, and Interspeech 2019 sleepiness challenge data.
- Integration of findings across speech, cough, breathing, and affective signal analysis to identify cross-modal biomarkers of respiratory and psychological health.
Experimental results
Research questions
- RQ1What audio and speech signal features can reliably indicate the presence of COVID-19 or its symptoms?
- RQ2How effective are AI-based models in detecting respiratory abnormalities such as cough, breathing rate, and OSA from speech and audio signals?
- RQ3To what extent can speech and audio processing support mental health monitoring during pandemic-related stress and isolation?
- RQ4What are the key challenges in deploying real-world systems for remote COVID-19 screening using audio modalities?
- RQ5How can speech and language processing be leveraged to improve public awareness and behavioral health during pandemics?
Key findings
- Cough sound analysis using deep learning models achieved up to 98.5% accuracy in distinguishing COVID-19 coughs from healthy and non-COVID-19 coughs.
- Speech-based detection of respiratory symptoms such as shortness of breath and vocal fatigue showed high sensitivity and specificity in controlled studies.
- A model using attention and autoencoder fusion achieved a Spearman’s correlation of 0.367 in predicting sleepiness from speech signals.
- Hierarchical Attention Transfer Networks achieved an RMSE of 5.66 and MAE of 4.28 on the PHQ-8 depression scale using the DAIC-WOZ dataset.
- The use of compact neural networks reduced model size by 95% without compromising recognition performance, enabling deployment on edge devices.
- Chatbots integrated with audio analysis showed potential for reducing psychological distress and improving health behavior, though design quality significantly impacted outcomes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.