Skip to main content
QUICK REVIEW

[Paper Review] Deep Net Features for Complex Emotion Recognition

Bhalaji Nagarajan, O.V. Ramana Murthy|arXiv (Cornell University)|Oct 31, 2018
Emotion and Mood RecognitionPsychology2 references3 citations
TL;DR

This paper proposes leveraging deep neural network features from pretrained models—AudioSet Net, VoxCeleb Net, and Deep Speech Net—for complex emotion recognition, particularly curiosity. By extracting and encoding high-level representations from deep layers of these networks into feature vectors, the approach achieves an F1 score of 0.85 on the EmoReact dataset, significantly outperforming the prior baseline of 0.69.

ABSTRACT

This paper investigates the influence of different acoustic features, audio-events based features and automatic speech translation based lexical features in complex emotion recognition such as curiosity. Pretrained networks, namely, AudioSet Net, VoxCeleb Net and Deep Speech Net trained extensively for different speech based applications are studied for this objective. Information from deep layers of these networks are considered as descriptors and encoded into feature vectors. Experimental results on the EmoReact dataset consisting of 8 complex emotions show the effectiveness, yielding highest F1 score of 0.85 as against the baseline of 0.69 in the literature.

Motivation & Objective

  • To investigate the effectiveness of diverse acoustic, audio-event, and lexical features in recognizing complex emotions such as curiosity.
  • To evaluate the utility of pretrained deep neural networks—AudioSet Net, VoxCeleb Net, and Deep Speech Net—for feature extraction in emotion recognition.
  • To explore whether high-level representations from deep layers of these models can serve as robust descriptors for complex emotion classification.
  • To improve performance on complex emotion recognition beyond existing baselines, particularly on the EmoReact dataset with 8 emotion classes.

Proposed method

  • Utilizing pretrained AudioSet Net, VoxCeleb Net, and Deep Speech Net models, each trained on large-scale speech and audio data for distinct tasks.
  • Extracting feature representations from the deeper, more abstract layers of these networks to capture high-level semantic patterns.
  • Encoding the extracted deep-layer features into compact, discriminative feature vectors for downstream classification.
  • Training a classifier on the combined deep features to recognize 8 complex emotions, including curiosity, on the EmoReact dataset.
  • Evaluating performance using standard metrics, particularly F1 score, to compare against prior state-of-the-art results.

Experimental results

Research questions

  • RQ1How do acoustic, audio-event, and lexical features derived from pretrained models contribute to complex emotion recognition?
  • RQ2To what extent can deep representations from AudioSet Net, VoxCeleb Net, and Deep Speech Net improve emotion recognition performance?
  • RQ3Can high-level features from pretrained networks generalize effectively to complex emotions like curiosity?
  • RQ4What performance gain is achievable over existing baselines using these deep net features?

Key findings

  • The proposed method achieved an F1 score of 0.85 on the EmoReact dataset, representing a significant improvement over the prior baseline of 0.69.
  • Deep features extracted from the higher layers of pretrained networks proved highly effective for capturing nuanced emotional states.
  • The combination of acoustic, audio-event, and automatic speech translation-based lexical features enhanced discriminative power for complex emotions.
  • Pretrained models fine-tuned for different speech applications still yield strong, transferable representations for emotion recognition.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.