[Paper Review] An Early Study on Intelligent Analysis of Speech under COVID-19: Severity, Sleep Quality, Fatigue, and Anxiety
This study presents the first empirical investigation into using speech analysis for COVID-19 symptom assessment, employing eGeMAPS and ComPARE acoustic features with SVMs to predict severity, sleep quality, fatigue, and anxiety in hospitalized patients. It achieved an average accuracy of 0.69 in classifying illness severity based on hospitalization duration, demonstrating the potential of audio-only models for low-cost, non-invasive screening.
The COVID-19 outbreak was announced as a global pandemic by the World Health Organisation in March 2020 and has affected a growing number of people in the past few weeks. In this context, advanced artificial intelligence techniques are brought to the fore in responding to fight against and reduce the impact of this global health crisis. In this study, we focus on developing some potential use-cases of intelligent speech analysis for COVID-19 diagnosed patients. In particular, by analysing speech recordings from these patients, we construct audio-only-based models to automatically categorise the health state of patients from four aspects, including the severity of illness, sleep quality, fatigue, and anxiety. For this purpose, two established acoustic feature sets and support vector machines are utilised. Our experiments show that an average accuracy of .69 obtained estimating the severity of illness, which is derived from the number of days in hospitalisation. We hope that this study can foster an extremely fast, low-cost, and convenient way to automatically detect the COVID-19 disease.
Motivation & Objective
- To develop audio-only models for automated assessment of COVID-19 patient health states, including illness severity, sleep quality, fatigue, and anxiety.
- To explore the feasibility of using speech signals as a non-invasive, low-cost diagnostic tool during the pandemic.
- To evaluate the performance of established acoustic feature sets (eGeMAPS and ComPARE) in predicting clinically relevant patient states from voice recordings.
- To establish a foundation for future intelligent speech analysis systems in pandemic response and remote patient monitoring.
Proposed method
- Collected 52 speech samples from hospitalized COVID-19 patients in Wuhan, China, during March 2020, with five neutral sentences per patient recorded via WeChat.
- Acquired self-reported ratings from patients on sleep quality, fatigue, and anxiety using a three-level scale (low, mid, high).
- Extracted two established acoustic feature sets—eGeMAPS and ComPARE—across full utterance segments for each recording.
- Trained separate Support Vector Machines (SVMs) for each of the four prediction tasks: severity, sleep quality, fatigue, and anxiety.
- Optimized SVM hyperparameters (C values) using cross-validation to maximize performance on each task.
- Evaluated models using unweighted average recall (UAR), accuracy (WAR), and F1-score, with chance level set at 0.33 for UAR.
Experimental results
Research questions
- RQ1Can speech signals alone be used to predict the severity of illness in hospitalized COVID-19 patients?
- RQ2To what extent can speech features predict self-reported mental and physical health states such as sleep quality, fatigue, and anxiety?
- RQ3How do different acoustic feature sets (eGeMAPS vs. ComPARE) compare in performance for these prediction tasks?
- RQ4What is the potential of SVM-based models in achieving reliable, non-invasive screening using only voice data?
Key findings
- The model achieved an average accuracy of 0.69 in predicting illness severity based on hospitalization duration, indicating strong potential for non-invasive screening.
- For sleep quality prediction, eGeMAPS achieved a UAR of 0.61 and WAR of 0.57, outperforming ComPARE (UAR: 0.49, WAR: 0.39).
- Fatigue prediction showed moderate performance, with eGeMAPS achieving a UAR of 0.46 and WAR of 0.50.
- Anxiety prediction reached a UAR of 0.56 with eGeMAPS and 0.53 with ComPARE, indicating detectable patterns in voice related to emotional state.
- The best-performing models used different C values for each task, with eGeMAPS generally outperforming ComPARE across all four tasks.
- Despite promising results, performance was limited by small dataset size and lack of control groups, suggesting room for improvement with larger, more diverse data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.