[Paper Review] Facial Expression Phoenix (FePh): An Annotated Sequenced Dataset for Facial and Emotion-Specified Expressions in Sign Language
This paper introduces FePh, the first publicly available annotated sequenced facial expression dataset in sign language, derived from the RWTH-PHOENIX-Weather 2014 dataset. It provides 3,000+ semi-blurry, dynamic facial images with diverse head poses, movements, and emotion labels across seven basic emotions and a 'None' class, enabling multi-modal sign language and gesture recognition research.
Facial expressions are important parts of both gesture and sign language recognition systems. Despite the recent advances in both fields, annotated facial expression datasets in the context of sign language are still scarce resources. In this manuscript, we introduce an annotated sequenced facial expression dataset in the context of sign language, comprising over $3000$ facial images extracted from the daily news and weather forecast of the public tv-station PHOENIX. Unlike the majority of currently existing facial expression datasets, FePh provides sequenced semi-blurry facial images with different head poses, orientations, and movements. In addition, in the majority of images, identities are mouthing the words, which makes the data more challenging. To annotate this dataset we consider primary, secondary, and tertiary dyads of seven basic emotions of "sad", "surprise", "fear", "angry", "neutral", "disgust", and "happy". We also considered the "None" class if the image's facial expression could not be described by any of the aforementioned emotions. Although we provide FePh as a facial expression dataset of signers in sign language, it has a wider application in gesture recognition and Human Computer Interaction (HCI) systems.
Motivation & Objective
- Address the scarcity of annotated facial expression datasets in sign language contexts, which hinders multi-modal sign language recognition research.
- Provide a vision-based, real-world dataset with dynamic facial expressions, head pose variations, and movement to better reflect natural sign language use.
- Enable research in facial expression recognition within sign language by offering detailed emotion annotations across seven basic emotions and a 'None' class.
- Support the development of multi-modal systems by combining facial expression data with existing handshape annotations from the RWTH-PHOENIX-Weather 2014 dataset.
- Facilitate broader applications in gesture recognition and Human-Computer Interaction (HCI) beyond sign language recognition.
Proposed method
- Extracted over 3,000 facial images from daily news and weather forecasts of the public TV station PHOENIX, focusing on signers with visible facial expressions.
- Annotated each image with primary, secondary, and tertiary emotion dyads using seven basic emotions: 'sad', 'surprise', 'fear', 'angry', 'neutral', 'disgust', 'happy', and a 'None' class for unclassifiable expressions.
- Preserved temporal sequences by maintaining the original video frame order, enabling analysis of dynamic facial expression transitions.
- Selected data from the RWTH-PHOENIX-Weather 2014 multisigner dataset to ensure consistency with an established benchmark in continuous sign language recognition.
- Applied a multi-level annotation scheme to capture nuanced emotional expressions, including co-occurring emotions (e.g., 'surprise_fear'), enhancing emotional granularity.
- Ensured diversity in head poses, orientations, and movements by selecting frames from natural broadcast videos, increasing dataset realism and challenge.
Experimental results
Research questions
- RQ1Can a large-scale, real-world facial expression dataset be effectively constructed from broadcast video data for sign language contexts?
- RQ2To what extent do facial expressions in sign language correlate with specific handshapes, and can such correlations be quantitatively measured?
- RQ3How do dynamic facial expressions with head movement and varying poses affect the performance of facial expression recognition models in sign language?
- RQ4Can the integration of facial expression and handshape data improve the accuracy of multi-modal sign language recognition systems?
- RQ5What is the distribution and frequency of facial expressions across different handshape classes in sign language, and does it indicate meaningful semantic or grammatical associations?
Key findings
- The FePh dataset contains over 3,000 facial images extracted from real broadcast videos, featuring diverse head poses, movements, and semi-blurry conditions.
- A significant correlation was found between handshape classes and specific facial expressions, with hand shape class '3' showing a 0.697 correlation with 'fear' and 0.383 with 'surprise'.
- The heatmap analysis confirmed that certain facial expressions, such as 'fear', are disproportionately expressed with specific handshapes, indicating grammatical or emotional significance.
- The dataset reveals that identical handshapes can be paired with different facial expressions, demonstrating the critical role of facial expressions in disambiguating meaning.
- The inclusion of 'None' class annotations highlights the complexity of real-world facial expression labeling, where not all expressions fit into basic emotion categories.
- The dataset enables future multi-modal learning by combining facial expression data with the existing RWTH-PHOENIX-Weather 2014 MS Handshapes dataset, offering a comprehensive resource for sign language recognition.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.