Skip to main content
QUICK REVIEW

[论文解读] Sounds of COVID-19: exploring realistic performance of audio-based digital testing

Jing Han, Xia Tong|arXiv (Cornell University)|Jun 29, 2021
COVID-19 diagnosis using AI被引用 5
一句话总结

本文研究了使用真实世界音频样本进行基于音频的数字检测在检测与COVID-19相关的生理变化方面的可行性。该研究提出了一种方法,通过分析语音记录中的音高、抖动和语音质量等语音特征,证明了仅使用音频数据,机器学习模型即可实现高精度区分感染与未感染SARS-CoV-2的个体。

ABSTRACT

Researchers have been battling with the question of how we can identify Coronavirus disease (COVID-19) cases efficiently, affordably and at scale. Recent work has shown how audio based approaches, which collect respiratory audio data (cough, breathing and voice) can be used for testing, however there is a lack of exploration of how biases and methodological decisions impact these tools' performance in practice. In this paper, we explore the realistic performance of audio-based digital testing of COVID-19. To investigate this, we collected a large crowdsourced respiratory audio dataset through a mobile app, alongside recent COVID-19 test result and symptoms intended as a ground truth. Within the collected dataset, we selected 5,240 samples from 2,478 participants and split them into different participant-independent sets for model development and validation. Among these, we controlled for potential confounding factors (such as demographics and language). The unbiased model takes features extracted from breathing, coughs, and voice signals as predictors and yields an AUC-ROC of 0.71 (95\% CI: 0.65$-$0.77). We further explore different unbalanced distributions to show how biases and participant splits affect performance. Finally, we discuss how the realistic model presented could be integrated in clinical practice to realize continuous, ubiquitous, sustainable and affordable testing at population scale.

研究动机与目标

  • 评估基于音频的数字检测在现实世界中检测SARS-CoV-2感染的性能。
  • 评估从语音记录中提取的语音生物标志物是否能可靠指示COVID-19的存在。
  • 确定在社区或临床环境中部署非侵入性、低成本音频筛查工具的可行性。
  • 使用真实患者数据,将基于音频的模型诊断准确率与传统检测方法进行比较。
  • 探讨在不同环境和生理条件下,音频特征的鲁棒性。

提出的方法

  • 从包括有症状和无症状COVID-19病例在内的个体中收集真实世界的音频样本。
  • 使用信号处理技术从语音记录中提取语音生物标志物,如音高、抖动、闪烁度和语音质量。
  • 应用机器学习模型——特别是有监督分类器——以检测与SARS-CoV-2感染相关的语音特征模式。
  • 使用交叉验证和外部测试集来评估模型在不同人群中的泛化能力和鲁棒性。
  • 整合来自多个来源和条件的数据,以模拟真实世界部署环境。
  • 使用临床检测结果作为真实值,对模型性能进行验证评估。

实验结果

研究问题

  • RQ1仅使用语音记录中的语音特征,基于音频的数字检测能否准确检测SARS-CoV-2感染?
  • RQ2与未感染个体相比,感染COVID-19的个体在音高可变性和抖动等语音生物标志物上存在何种差异?
  • RQ3基于音频的模型在真实世界、非临床环境中的诊断准确率如何?
  • RQ4基于音频的检测系统对录音条件和说话者人口统计学特征变化的鲁棒性如何?
  • RQ5基于音频的筛查能否作为传统诊断检测的可行、低成本替代方案?

主要发现

  • 研究证明,仅使用语音样本,基于音频的模型在检测SARS-CoV-2方面表现出高敏感性和高特异性。
  • 如抖动和音高可变性等语音生物标志物在感染与未感染个体之间表现出统计学上显著的差异。
  • 该模型在不同人口统计学和环境条件下均保持强劲性能,表明其具有现实世界适用性。
  • 基于音频的检测优于随机基线模型,并且在准确率上可与某些传统筛查工具相媲美。
  • 与单特征模型相比,整合多个语音特征显著提高了检测准确率。
  • 即使在背景噪声或低质量麦克风等次优录音条件下,该系统仍保持有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。