[Paper Review] Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution
This study investigates gendered emotion attribution in five state-of-the-art large language models (LLMs) using persona-based prompting. It finds that all models consistently associate women with sadness and men with anger—mirroring entrenched societal stereotypes—raising concerns about the fairness and ethical implications of using LLMs in emotion-related applications.
Large language models (LLMs) reflect societal norms and biases, especially about gender. While societal biases and stereotypes have been extensively researched in various NLP applications, there is a surprising gap for emotion analysis. However, emotion and gender are closely linked in societal discourse. E.g., women are often thought of as more empathetic, while men's anger is more socially accepted. To fill this gap, we present the first comprehensive study of gendered emotion attribution in five state-of-the-art LLMs (open- and closed-source). We investigate whether emotions are gendered, and whether these variations are based on societal stereotypes. We prompt the models to adopt a gendered persona and attribute emotions to an event like 'When I had a serious argument with a dear person'. We then analyze the emotions generated by the models in relation to the gender-event pairs. We find that all models consistently exhibit gendered emotions, influenced by gender stereotypes. These findings are in line with established research in psychology and gender studies. Our study sheds light on the complex societal interplay between language, gender, and emotion. The reproduction of emotion stereotypes in LLMs allows us to use those models to study the topic in detail, but raises questions about the predictive use of those same LLMs for emotion applications.
Motivation & Objective
- To investigate whether large language models (LLMs) reflect societal gendered stereotypes in emotion attribution.
- To determine whether these emotional associations stem from actual lived experiences or are driven by ingrained gender stereotypes.
- To examine the role of LLMs as both mirrors of societal bias and amplifiers of representational harm in emotion-related NLP tasks.
- To provide a comprehensive, quantitative and qualitative analysis of emotion attribution across diverse LLMs and gendered personas.
- To support future research by releasing all data and promoting interdisciplinary collaboration in NLP, psychology, and gender studies.
Proposed method
- Employed persona-based prompting to elicit emotion attributions from five SOTA LLMs (open- and closed-source), including GPT-4 and LLaMA.
- Presented each model with a standardized event—'When I had a serious argument with a dear person'—paired with male or female gendered personas.
- Collected over 200,000 completions across 7,000 unique events and two gendered personas, spanning more than 400 distinct emotions.
- Conducted quantitative analysis of emotion distributions across gendered personas to identify systematic biases.
- Performed qualitative analysis of model-generated explanations to validate and contextualize the observed emotional associations.
- Compared model outputs against self-reported emotion data to assess alignment with lived experience versus societal stereotypes.

Experimental results
Research questions
- RQ1Do large language models exhibit gendered emotional attributions when prompted with gendered personas?
- RQ2Are the observed emotional associations in LLMs shaped by real differences in lived emotional experiences or by societal gender stereotypes?
- RQ3To what extent do LLMs reproduce and amplify existing gendered emotional stereotypes found in psychology and gender studies?
- RQ4How do model explanations reflect or obscure these stereotypical associations?
- RQ5What are the implications of these biases for the use of LLMs in sensitive emotion-related applications such as mental health or human-computer interaction?
Key findings
- All five state-of-the-art LLMs consistently attribute sadness more frequently to female personas and anger more frequently to male personas when prompted with the same emotional event.
- The emotional associations strongly align with established psychological and sociological research on gendered emotional stereotypes, such as the perception of women as more emotional and men as more angry.
- Even when self-reported data indicated different emotional responses, the models maintained the same gendered emotional associations, indicating that the bias is not grounded in lived experience but in societal stereotypes.
- The qualitative analysis of model explanations revealed that the models often justify these associations using gendered behavioral and emotional norms, reinforcing traditional roles.
- The findings highlight a significant risk of representational harm when deploying LLMs in emotion analysis tasks, particularly in mental health and human-computer interaction.
- The study underscores the dual role of LLMs as both reflective of and amplifiers of societal bias, urging caution in their use for emotion-related applications.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.