Skip to main content
QUICK REVIEW

[Paper Review] Linguistic and Gender Variation in Speech Emotion Recognition using Spectral Features

Zachary Dair, Ryan Donovan|arXiv (Cornell University)|Dec 17, 2021
Music and Audio Processing4 citations
TL;DR

This study investigates how gender and linguistic differences affect speech emotion recognition (SER) using spectral features such as Mel-frequency Cepstral Coefficients (MFCC), Melspectrograms, and Spectral Contrast. A convolutional neural network (CNN) was trained on English, German, and Italian datasets, revealing that emotion classification accuracy is high in mono- and cross-lingual settings but significantly lower in multi-lingual data, with notable performance differences between male and female speakers, especially for high-energy (e.g., Anger, Joy) and low-energy (e.g., Disgust, Fear) emotions.

ABSTRACT

This work explores the effect of gender and linguistic-based vocal variations on the accuracy of emotive expression classification. Emotive expressions are considered from the perspective of spectral features in speech (Mel-frequency Cepstral Coefficient, Melspectrogram, Spectral Contrast). Emotions are considered from the perspective of Basic Emotion Theory. A convolutional neural network is utilised to classify emotive expressions in emotive audio datasets in English, German, and Italian. Vocal variations for spectral features assessed by (i) a comparative analysis identifying suitable spectral features, (ii) the classification performance for mono, multi and cross-lingual emotive data and (iii) an empirical evaluation of a machine learning model to assess the effects of gender and linguistic variation on classification accuracy. The results showed that spectral features provide a potential avenue for increasing emotive expression classification. Additionally, the accuracy of emotive expression classification was high within mono and cross-lingual emotive data, but poor in multi-lingual data. Similarly, there were differences in classification accuracy between gender populations. These results demonstrate the importance of accounting for population differences to enable accurate speech emotion recognition.

Motivation & Objective

  • To examine the impact of gender and linguistic variation on speech emotion recognition (SER) using spectral features.
  • To evaluate the performance of a CNN-based model across mono-, multi-, and cross-lingual speech emotion classification.
  • To identify which spectral features (MFCC, Melspectrogram, Spectral Contrast) are most effective for emotion classification across diverse populations.
  • To assess how vocal differences related to gender and language affect the generalizability and accuracy of SER systems.
  • To highlight the need for standardized data collection and normalization techniques to improve cross-linguistic and cross-gender SER performance.

Proposed method

  • Spectral features including Mel-frequency Cepstral Coefficients (MFCC), Melspectrograms, and Spectral Contrast were extracted from audio data in English, German, and Italian.
  • A convolutional neural network (CNN) was trained and evaluated for emotion classification using the extracted spectral features.
  • The model was tested across three settings: mono-lingual (same language), multi-lingual (multiple languages in training and test), and cross-lingual (training on one language, testing on another).
  • Comparative analysis identified the most effective spectral features for emotion classification across different gender and linguistic groups.
  • Empirical evaluation assessed the influence of gender and language on classification accuracy using performance metrics across the six basic emotions (Anger, Disgust, Fear, Joy, Sadness, Surprise).
  • Statistical evaluation was conducted to determine whether differences in performance were due to inherent population variations or data sampling limitations.

Experimental results

Research questions

  • RQ1How do gender differences in vocal characteristics affect the accuracy of speech emotion recognition using spectral features?
  • RQ2To what extent does linguistic variation across English, German, and Italian impact the performance of CNN-based emotion classification models?
  • RQ3Which spectral features (MFCC, Melspectrogram, Spectral Contrast) are most effective for emotion classification across diverse gender and linguistic populations?
  • RQ4Why does multi-lingual SER performance degrade compared to mono- and cross-lingual setups, and what role do recording variability and normalization play?
  • RQ5How do signal energy and amplitude characteristics influence emotion detection accuracy in male versus female speakers across different languages?

Key findings

  • The CNN model achieved high emotion classification accuracy in mono- and cross-lingual settings, but performance dropped significantly in multi-lingual setups due to linguistic and recording variability.
  • High-energy emotions such as Anger and Joy were more accurately detected in German-speaking males, likely due to higher amplitude and harsher vocal characteristics in this group.
  • Low-energy emotions such as Disgust and Fear were more accurately classified from female voices across all languages, possibly due to lower amplitude and softer vocal articulation.
  • Significant differences in spectral representation between male and female speakers led to gender-specific performance disparities, with combined-gender models underperforming compared to gender-specific models.
  • The study identified that spectral features related to amplitude and energy are critical for accurate emotion recognition, especially when accounting for gender and linguistic variation.
  • Small sample sizes (N = 40) and unequal speaker distribution across linguistic groups may have exaggerated observed differences, limiting generalizability and suggesting a need for larger, balanced datasets in future work.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.