[Paper Review] Linguistic and Gender Variation in Speech Emotion Recognition using Spectral Features
This study investigates how gender and linguistic differences affect speech emotion recognition (SER) using spectral features such as Mel-frequency Cepstral Coefficients (MFCC), Melspectrograms, and Spectral Contrast. A convolutional neural network (CNN) was trained on English, German, and Italian datasets, revealing that emotion classification accuracy is high in mono- and cross-lingual settings but significantly lower in multi-lingual data, with notable performance differences between male and female speakers, especially for high-energy (e.g., Anger, Joy) and low-energy (e.g., Disgust, Fear) emotions.
This work explores the effect of gender and linguistic-based vocal variations on the accuracy of emotive expression classification. Emotive expressions are considered from the perspective of spectral features in speech (Mel-frequency Cepstral Coefficient, Melspectrogram, Spectral Contrast). Emotions are considered from the perspective of Basic Emotion Theory. A convolutional neural network is utilised to classify emotive expressions in emotive audio datasets in English, German, and Italian. Vocal variations for spectral features assessed by (i) a comparative analysis identifying suitable spectral features, (ii) the classification performance for mono, multi and cross-lingual emotive data and (iii) an empirical evaluation of a machine learning model to assess the effects of gender and linguistic variation on classification accuracy. The results showed that spectral features provide a potential avenue for increasing emotive expression classification. Additionally, the accuracy of emotive expression classification was high within mono and cross-lingual emotive data, but poor in multi-lingual data. Similarly, there were differences in classification accuracy between gender populations. These results demonstrate the importance of accounting for population differences to enable accurate speech emotion recognition.
Motivation & Objective
- To examine the impact of gender and linguistic variation on speech emotion recognition (SER) using spectral features.
- To evaluate the performance of a CNN-based model across mono-, multi-, and cross-lingual speech emotion classification.
- To identify which spectral features (MFCC, Melspectrogram, Spectral Contrast) are most effective for emotion classification across diverse populations.
- To assess how vocal differences related to gender and language affect the generalizability and accuracy of SER systems.
- To highlight the need for standardized data collection and normalization techniques to improve cross-linguistic and cross-gender SER performance.
Proposed method
- Spectral features including Mel-frequency Cepstral Coefficients (MFCC), Melspectrograms, and Spectral Contrast were extracted from audio data in English, German, and Italian.
- A convolutional neural network (CNN) was trained and evaluated for emotion classification using the extracted spectral features.
- The model was tested across three settings: mono-lingual (same language), multi-lingual (multiple languages in training and test), and cross-lingual (training on one language, testing on another).
- Comparative analysis identified the most effective spectral features for emotion classification across different gender and linguistic groups.
- Empirical evaluation assessed the influence of gender and language on classification accuracy using performance metrics across the six basic emotions (Anger, Disgust, Fear, Joy, Sadness, Surprise).
- Statistical evaluation was conducted to determine whether differences in performance were due to inherent population variations or data sampling limitations.
Experimental results
Research questions
- RQ1How do gender differences in vocal characteristics affect the accuracy of speech emotion recognition using spectral features?
- RQ2To what extent does linguistic variation across English, German, and Italian impact the performance of CNN-based emotion classification models?
- RQ3Which spectral features (MFCC, Melspectrogram, Spectral Contrast) are most effective for emotion classification across diverse gender and linguistic populations?
- RQ4Why does multi-lingual SER performance degrade compared to mono- and cross-lingual setups, and what role do recording variability and normalization play?
- RQ5How do signal energy and amplitude characteristics influence emotion detection accuracy in male versus female speakers across different languages?
Key findings
- The CNN model achieved high emotion classification accuracy in mono- and cross-lingual settings, but performance dropped significantly in multi-lingual setups due to linguistic and recording variability.
- High-energy emotions such as Anger and Joy were more accurately detected in German-speaking males, likely due to higher amplitude and harsher vocal characteristics in this group.
- Low-energy emotions such as Disgust and Fear were more accurately classified from female voices across all languages, possibly due to lower amplitude and softer vocal articulation.
- Significant differences in spectral representation between male and female speakers led to gender-specific performance disparities, with combined-gender models underperforming compared to gender-specific models.
- The study identified that spectral features related to amplitude and energy are critical for accurate emotion recognition, especially when accounting for gender and linguistic variation.
- Small sample sizes (N = 40) and unequal speaker distribution across linguistic groups may have exaggerated observed differences, limiting generalizability and suggesting a need for larger, balanced datasets in future work.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.