Skip to main content
QUICK REVIEW

[论文解读] A Crowdsourcing Extension of the ITU-T Recommendation P.835 with Validation.

Babak Naderi, Ross Cutler|arXiv (Cornell University)|Oct 25, 2020
Speech and Audio Processing参考文献 11被引用 12
一句话总结

本文提出了一种开源的、自动化的众包扩展方法,用于ITU-T Rec. P.835语音质量评估,其设计符合P.808指南。与实验室环境下的P.835评估相比,该方法在有效性方面表现优异(平均PCC = 0.961),可重复性极高(PCC = 1.00),并已证明在基准测试深度噪声抑制模型方面具有实用性。

ABSTRACT

The quality of the speech communication systems, which include noise suppression algorithms, are typically evaluated in laboratory experiments according to the ITU-T Rec. P.835. In this paper, we introduce an open-source implementation of the ITU-T Rec. P.835 for the crowdsourcing approach following the ITU-T Rec. P.808 on crowdsourcing recommendations. The implementation is an extension of the P.808 Toolkit and is highly automated to avoid operational errors. To assess our evaluation method's validity, we compared the Mean Opinion Scores (MOS), calculate using ratings collected with our implementation, and the MOS values from a standard laboratory experiment conducted according to the ITU-T Rec, P.835. Results show a high validity in all three scales (average PCC = 0.961). Results of a round-robin test showed that our implementation is a highly reproducible evaluation method (PCC=1.00). Finally, we investigated the performance of five models deep noise suppression models using our P.835 implementation and show what insights can be learned.

研究动机与目标

  • 开发一种可扩展的、基于众包的ITU-T Rec. P.835实现方法,用于语音质量评估。
  • 通过与标准实验室环境下的P.835评估进行对比,验证众包MOS评分的可靠性。
  • 通过众包框架在多次测试轮次中确保方法论的可重复性。
  • 将所提出的方法应用于评估和比较五种深度噪声抑制模型的性能。

提出的方法

  • 扩展ITU-T Rec. P.808工具包,以支持使用P.835 MOS方法进行众包语音质量测试。
  • 自动化整个评估流程,以最大限度减少操作错误并确保一致性。
  • 通过标准化的众包平台收集参与者的平均意见得分(MOS)。
  • 使用皮尔逊相关系数(PCC)比较众包MOS与实验室P.835 MOS值。
  • 采用轮换测试协议,评估在多次运行中方法论的可重复性。
  • 采用五种深度噪声抑制模型作为测试案例,以评估该框架的实际应用价值。

实验结果

研究问题

  • RQ1与标准实验室环境下的P.835评估相比,众包MOS评估在有效性方面表现如何?
  • RQ2所提出的众包实现方法在多次重复测试运行中具有多大程度的可重复性?
  • RQ3所提出的框架能否可靠地区分不同深度噪声抑制模型的性能差异?
  • RQ4通过众包获得的MOS评分与传统实验室实验结果之间的相关性如何?

主要发现

  • 众包MOS评分表现出高度有效性,与实验室P.835评估相比,平均皮尔逊相关系数(PCC)达到0.961。
  • 该方法展现出完美的可重复性,在多次评估轮次的轮换测试中,PCC达到1.00。
  • 该框架成功实现了对五种深度噪声抑制模型的比较,揭示了与主观感知预期一致的性能差异。
  • 自动化实现显著减少了操作错误,从而提高了语音质量评估的可靠性与可扩展性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。