[Paper Review] Addressing Racial Bias in Facial Emotion Recognition
This study investigates racial bias in facial emotion recognition (FER) by simulating diverse racial compositions in training data from AffectNet and CAFE datasets. Using sub-sampled training sets with varying racial distributions, it finds that racial balance improves fairness and performance—especially in smaller, posed datasets—yet fails to eliminate bias in larger, more diverse datasets, indicating that compositional balance alone is insufficient to address deeper sources of bias such as annotation and model estimation errors.
Fairness in deep learning models trained with high-dimensional inputs and subjective labels remains a complex and understudied area. Facial emotion recognition, a domain where datasets are often racially imbalanced, can lead to models that yield disparate outcomes across racial groups. This study focuses on analyzing racial bias by sub-sampling training sets with varied racial distributions and assessing test performance across these simulations. Our findings indicate that smaller datasets with posed faces improve on both fairness and performance metrics as the simulations approach racial balance. Notably, the F1-score increases by $27.2\%$ points, and demographic parity increases by $15.7\%$ points on average across the simulations. However, in larger datasets with greater facial variation, fairness metrics generally remain constant, suggesting that racial balance by itself is insufficient to achieve parity in test performance across different racial groups.
Motivation & Objective
- To investigate how racial composition in training data affects fairness and performance in facial emotion recognition (FER) models.
- To assess whether achieving racial balance in training data mitigates disparities in FER performance across different racial groups.
- To identify persistent sources of bias beyond data composition, such as annotation and race estimation errors.
- To evaluate the limitations of compositional bias mitigation in real-world, high-dimensional, subjective-label datasets like AffectNet and CAFE.
Proposed method
- Sub-sampled training sets from AffectNet and CAFE datasets with varying racial proportions to simulate different training distributions.
- Trained FER models on each sub-sampled dataset and evaluated performance using accuracy, F1-score, and fairness metrics (demographic parity, equalized odds).
- Used the FairFace model to estimate race for images in AffectNet and CAFE to enable race-specific evaluation and simulation.
- Compared model performance across simulations with different racial compositions, focusing on race-specific F1-scores and fairness metrics.
- Analyzed annotation bias by examining labeler demographics and label agreement rates, particularly in AffectNet’s limited-labeler setup.
- Proposed excluding less accurately estimated racial groups (e.g., Middle Eastern, South Asian) in future simulations to isolate estimator bias.

Experimental results
Research questions
- RQ1How does the racial composition of training data affect race-specific F1-scores and fairness metrics in FER models?
- RQ2To what extent does achieving racial balance in training data improve fairness and performance across different racial groups?
- RQ3Why do fairness metrics remain suboptimal even when training data is racially balanced?
- RQ4What role do annotation biases and race estimation errors play in perpetuating disparities in FER models?
Key findings
- In smaller, posed datasets (CAFE), racial balance improved fairness: demographic parity increased by 15.7 percentage points and F1-score by 27.2 percentage points on average.
- In larger, more diverse datasets (AffectNet), fairness metrics showed minimal improvement despite racial balancing, indicating that compositional balance alone is insufficient.
- The F1-score for 'angry' expressions improved only when East Asian faces were over-sampled, suggesting non-uniform effects of data composition.
- Race estimation errors from the FairFace model were evident, particularly for Middle Eastern and South Asian individuals, potentially distorting simulation results.
- Annotation bias in AffectNet—due to limited, non-diverse labelers—likely contributes to persistent disparities, as agreement rates varied from 50.8% (neutral) to 79.6% (happy).
- Even with balanced training data, fairness metrics like equalized odds did not consistently improve, suggesting unaccounted biases such as algorithmic or feature-level bias persist.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.