Skip to main content
QUICK REVIEW

[Paper Review] Context-Dependent Models for Predicting and Characterizing Facial Expressiveness

Victoria Lin, Jeffrey M. Girard|arXiv (Cornell University)|Dec 10, 2019
Emotion and Mood Recognition17 references4 citations
TL;DR

This paper introduces an extended version of the BP4D+ dataset with human-annotated expressiveness ratings and proposes context-dependent models to predict momentary facial expressiveness from visual data. Using statistical and deep learning models, it demonstrates that performance improves significantly when modeling expressiveness separately in emotional contexts (e.g., startle, pain, disgust), achieving near-human correlation with ground truth, while revealing context-specific visual signals such as action unit intensity and facial displacement that drive expressiveness.

ABSTRACT

In recent years, extensive research has emerged in affective computing on topics like automatic emotion recognition and determining the signals that characterize individual emotions. Much less studied, however, is expressiveness, or the extent to which someone shows any feeling or emotion. Expressiveness is related to personality and mental health and plays a crucial role in social interaction. As such, the ability to automatically detect or predict expressiveness can facilitate significant advancements in areas ranging from psychiatric care to artificial social intelligence. Motivated by these potential applications, we present an extension of the BP4D+ dataset with human ratings of expressiveness and develop methods for (1) automatically predicting expressiveness from visual data and (2) defining relationships between interpretable visual signals and expressiveness. In addition, we study the emotional context in which expressiveness occurs and hypothesize that different sets of signals are indicative of expressiveness in different contexts (e.g., in response to surprise or in response to pain). Analysis of our statistical models confirms our hypothesis. Consequently, by looking at expressiveness separately in distinct emotional contexts, our predictive models show significant improvements over baselines and achieve comparable results to human performance in terms of correlation with the ground truth.

Motivation & Objective

  • To develop automatic methods for predicting momentary facial expressiveness from visual data, which is critical for applications in artificial social intelligence and mental health assessment.
  • To identify and characterize interpretable visual signals—such as facial movements, gestures, and body posture—that underlie expressiveness in different emotional contexts.
  • To test the hypothesis that the set of visual signals indicative of expressiveness varies depending on the emotional context (e.g., startle vs. pain vs. disgust).
  • To create a unified, latent-variable expressiveness score from multi-aspect human ratings (response strength, emotion intensity, movement) to enable robust modeling.
  • To improve predictive performance by training context-specific models rather than a single general model across all emotional states.

Proposed method

  • Extended the BP4D+ dataset with human annotations for response strength, emotion intensity, and body/facial movement, then derived a latent variable-based expressiveness score using probabilistic modeling.
  • Trained both deep learning and interpretable statistical models (including ElasticNet) on visual features extracted via OpenFace, including facial landmarks, action units, and motion trajectories.
  • Designed context-specific models by training on data segmented by emotional context (startle, pain, disgust), enabling modeling of context-dependent signal contributions.
  • Used feature weights from interpretable linear models (ElasticNet) to analyze and visualize the relationship between specific visual signals and expressiveness in each context.
  • Evaluated models using correlation and NRMSE metrics against ground truth and human baseline performance, with statistical significance testing.
  • Applied latent variable modeling to integrate multi-aspect human ratings into a single, unified expressiveness score for training and evaluation.

Experimental results

Research questions

  • RQ1Does modeling expressiveness separately in distinct emotional contexts (e.g., startle, pain, disgust) lead to improved predictive performance compared to a single unified model?
  • RQ2Which specific visual signals—such as facial action units, motion velocity, or displacement—are most predictive of expressiveness in different emotional contexts?
  • RQ3To what extent do the same visual signals contribute to expressiveness across different emotional states, and how do they differ?
  • RQ4Can context-specific models achieve performance comparable to human-level correlation with ground truth in expressiveness prediction?
  • RQ5How do interpretable model weights reveal the psychological plausibility of identified visual signals in relation to known emotional responses?

Key findings

  • Context-specific models significantly outperformed general models and baselines, achieving a correlation of 0.84 with ground truth, approaching the human baseline correlation of 0.86.
  • The ElasticNet model achieved a correlation of 0.84 with ground truth, outperforming all baselines, though it still showed significantly higher NRMSE and lower correlation than the human baseline.
  • Across all contexts, facial motion features such as action unit count, action unit intensity, and point displacement were the most predictive signals of expressiveness.
  • In the startle context, high expressiveness was associated with increased point displacement and velocity, reduced head displacement, and higher action unit count—consistent with a reflexive startle response.
  • In the pain context, high expressiveness was linked to higher action unit intensity and count, but lower point velocity, suggesting a regulatory or controlled response rather than increased motion.
  • In the disgust context, high expressiveness was associated with increased action unit intensity, point displacement, and head displacement—consistent with recoiling from a distasteful stimulus.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.