Skip to main content
QUICK REVIEW

[Paper Review] InQSS: a speech intelligibility assessment model using a multi-task learning network.

Yu‐Wen Chen, Yu Tsao|arXiv (Cornell University)|Nov 4, 2021
Speech and Audio Processing27 references4 citations
TL;DR

InQSS is a multi-task learning speech intelligibility assessment model that leverages spectrogram and scattering coefficients as input features, jointly predicting both intelligibility and quality scores. Experimental results demonstrate that scattering coefficients and quality scores significantly improve intelligibility prediction, with the proposed TMHINT-QI dataset released for Chinese speech evaluation.

ABSTRACT

Speech intelligibility assessment models are essential tools for researchers to evaluate and improve speech processing models. In this study, we propose InQSS, a speech intelligibility assessment model that uses both spectrogram and scattering coefficients as input features. In addition, InQSS uses a multi-task learning network in which quality scores can guide the training of the speech intelligibility assessment. The resulting model can predict not only the intelligibility scores but also the quality scores of a speech. The experimental results confirm that the scattering coefficients and quality scores are informative for intelligibility. Moreover, we released TMHINT-QI, which is a Chinese speech dataset that records the quality and intelligibility scores of clean, noisy, and enhanced speech.

Motivation & Objective

  • To develop a robust speech intelligibility assessment model that improves prediction accuracy by integrating multiple signal representations.
  • To investigate the effectiveness of scattering coefficients as informative features for speech intelligibility modeling.
  • To explore the benefit of multi-task learning by jointly training on intelligibility and quality scores.
  • To create and release a new Chinese speech dataset (TMHINT-QI) with annotated quality and intelligibility scores for clean, noisy, and enhanced speech.
  • To provide a unified framework that predicts both intelligibility and quality scores from speech signals.

Proposed method

  • InQSS uses spectrogram and scattering coefficients as dual-input features to capture both time-frequency and invariant signal characteristics.
  • The model employs a multi-task learning architecture where shared and task-specific branches predict intelligibility and quality scores simultaneously.
  • Scattering coefficients are computed using a wavelet-based transform to extract stable, low-level features sensitive to speech degradation.
  • The multi-task loss function combines intelligibility and quality prediction losses to jointly optimize both objectives.
  • The network is trained end-to-end on the newly released TMHINT-QI dataset containing diverse speech conditions.
  • Feature-level fusion is applied before the final prediction heads to enable cross-task feature interaction.

Experimental results

Research questions

  • RQ1Can scattering coefficients improve the performance of speech intelligibility assessment models compared to spectrograms alone?
  • RQ2To what extent do quality scores enhance the prediction of speech intelligibility in a multi-task learning framework?
  • RQ3How effective is the joint learning of intelligibility and quality prediction in improving model generalization?
  • RQ4What is the contribution of different input features (spectrogram vs. scattering) to the final intelligibility prediction?
  • RQ5How well does the proposed model generalize across clean, noisy, and enhanced speech conditions?

Key findings

  • The integration of scattering coefficients significantly improves intelligibility prediction performance compared to models using spectrograms alone.
  • Joint training with quality scores enhances the model's ability to generalize across diverse speech degradation types.
  • The multi-task learning framework leads to more stable and accurate predictions for both intelligibility and quality scores.
  • The experimental results confirm that both scattering coefficients and quality scores are informative features for intelligibility assessment.
  • The released TMHINT-QI dataset provides a valuable benchmark for evaluating speech intelligibility and quality in Chinese speech processing.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.