[Paper Review] Analyzing the Affect of a Group of People Using Multi-modal Framework
This paper proposes a multi-modal framework for group-level emotion recognition (GER) by fusing face, upper-body, and scene-level cues using a novel information aggregation method. The approach achieves state-of-the-art performance on two challenging databases, with a mean absolute error (MAE) of 0.4835 on HAPPEI and 66.67% accuracy on GAFF, demonstrating that multi-modal fusion significantly improves robustness and accuracy in real-world, unconstrained environments.
Millions of images on the web enable us to explore images from social events such as a family party, thus it is of interest to understand and model the affect exhibited by a group of people in images. But analysis of the affect expressed by multiple people is challenging due to varied indoor and outdoor settings, and interactions taking place between various numbers of people. A few existing works on Group-level Emotion Recognition (GER) have investigated on face-level information. Due to the challenging environments, face may not provide enough information to GER. Relatively few studies have investigated multi-modal GER. Therefore, we propose a novel multi-modal approach based on a new feature description for understanding emotional state of a group of people in an image. In this paper, we firstly exploit three kinds of rich information containing face, upperbody and scene in a group-level image. Furthermore, in order to integrate multiple person's information in a group-level image, we propose an information aggregation method to generate three features for face, upperbody and scene, respectively. We fuse face, upperbody and scene information for robustness of GER against the challenging environments. Intensive experiments are performed on two challenging group-level emotion databases to investigate the role of face, upperbody and scene as well as multi-modal framework. Experimental results demonstrate that our framework achieves very promising performance for GER.
Motivation & Objective
- To address the challenge of group-level emotion recognition in unconstrained, real-world environments where face detection often fails due to occlusions, lighting, and pose variations.
- To investigate the complementary role of upper-body and scene-level cues in enhancing emotion recognition when face-level information is unreliable.
- To develop a robust information aggregation method that encodes collective emotional states from multiple individuals in a group image.
- To design and evaluate a multi-modal fusion framework that integrates face, upper-body, and scene features for improved GER performance.
- To validate the framework on two challenging databases—HAPPEI and GAFF—under both regression (happiness intensity) and classification (positive/neutral/negative) tasks.
Proposed method
- The framework extracts three types of features: face-level (using facial expressions), upper-body-level (using pose and posture), and scene-level (using contextual environmental cues).
- An information aggregation method is proposed to encode collective emotional states from multiple individuals, transforming individual-level features into group-level representations.
- A multi-modal fusion strategy combines face, upper-body, and scene features using multiple kernel learning (MKL), with a revisited localized MKL approach for improved generalization.
- The system employs a hybrid bottom-up (individual attributes) and top-down (scene and group structure) modeling approach inspired by social psychology principles.
- The framework is evaluated using both regression (mean absolute error on HAPPEI) and classification (recognition accuracy on GAFF) metrics.
- Parameter tuning is performed on the HAPPEI database and applied to the GAFF database, ensuring consistency in model configuration across datasets.
Experimental results
Research questions
- RQ1Can upper-body and scene-level features improve group-level emotion recognition when face detection fails due to real-world challenges?
- RQ2How does multi-modal fusion of face, upper-body, and scene information enhance robustness and accuracy in unconstrained group emotion recognition?
- RQ3What is the relative contribution of face, upper-body, and scene modalities to group-level affect prediction in diverse social settings?
- RQ4Does the proposed information aggregation method effectively encode collective emotional states from multiple individuals in a group image?
- RQ5How does the proposed multi-modal framework compare to existing state-of-the-art methods like GEM and CCRF-based models on benchmark datasets?
Key findings
- The proposed multi-modal framework achieves a mean absolute error (MAE) of 0.4835 on the HAPPEI database, outperforming all single-modality baselines and existing GEM models.
- Face-only performance (MAE = 0.5187) is comparable to state-of-the-art GEM models, but multi-modal fusion reduces MAE by 0.0457, demonstrating significant improvement.
- The inclusion of upper-body and scene features contributes meaningfully to performance, with scene alone achieving 0.7273 MAE in the HI kernel, indicating its value in context-aware emotion inference.
- On the GAFF database, the multi-modal system achieves 66.67% recognition accuracy, outperforming face-only (58.33%) and upper-body-only (46.57%) baselines, and approaching the performance of a more complex feature combination (67.64%).
- The proposed information aggregation method effectively captures group-level affect by integrating individual-level cues into a unified representation, enhancing robustness in challenging environments.
- The results confirm that multi-modal fusion is essential for robust group-level emotion recognition, especially when face-level data is unreliable due to occlusion, pose, or lighting variations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.