[Paper Review] A Friendly Face: Do Text-to-Image Systems Rely on Stereotypes when the Input is Under-Specified?
The paper investigates whether under-specified prompts in text-to-image models (DALL-E 2, Midjourney, Stable Diffusion) induce demographic stereotypes in generated faces using the ABC model of social cognition.
As text-to-image systems continue to grow in popularity with the general public, questions have arisen about bias and diversity in the generated images. Here, we investigate properties of images generated in response to prompts which are visually under-specified, but contain salient social attributes (e.g., 'a portrait of a threatening person' versus 'a portrait of a friendly person'). Grounding our work in social cognition theory, we find that in many cases, images contain similar demographic biases to those reported in the stereotype literature. However, trends are inconsistent across different models and further investigation is warranted.
Motivation & Objective
- Motivate examination of bias and diversity in text-to-image outputs when prompts lack explicit demographic detail.
- Ground analysis in the ABC Model of social cognition (Agency, Beliefs, Communion) to map prompts to demographic inferences.
- Assess whether three popular models reproduce stereotypical demographics for traits derived from the ABC dimensions.
- Evaluate inter-annotator reliability and ambiguity handling in annotating perceived demographics of generated images.
- Discuss implications for model design and potential bias mitigation strategies at inference and training stages.
Proposed method
- Use prompts of the form 'portrait of a <adjective> person' to generate images across three models (DALL-E 2, Midjourney, Stable Diffusion).
- Adapt prompts with ABC traits (high/low Agency, Beliefs, Communion) and generate multiple images per trait pole to assess bias.
- Annotate generated images for perceived gender, skin colour, and age via three annotators, then average annotations to obtain demographic scores.
- Analyze inter-annotator agreement with Cohen's Kappa and report across models.
- Examine ambiguity (AIAO) vs diversity (AIDO) strategies and report the proportion of images with ambiguous cues.
Experimental results
Research questions
- RQ1Do text-to-image models reproduce demographic biases when prompted with under-specified social traits?
- RQ2How do different models (DALL-E 2, Midjourney, Stable Diffusion) differ in gender, skin colour, and age tendencies for ABC-model prompts?
- RQ3Are there observable intersectional biases (e.g., poor associated with darker-skinned males) across models?
- RQ4What is the role of under-specification, ambiguity, and prompt engineering in shaping output bias?
- RQ5What implications do these biases have for debiasing approaches and model design?
Key findings
- Models show idiosyncratic biases along ABC dimensions; not all traits yield stereotypical images, but patterns exist per model.
- Midjourney and Stable Diffusion tend to produce lighter-skinned and younger subjects; DALL-E shows a general male over-representation in the baseline.
- High-Agency prompts often bias toward male gender and high-Communion toward female appearance in some models; low-Communion biases toward male gender are observed in others.
- Progressive beliefs often associate with lighter skin in Midjourney; DALL-E and Stable Diffusion show lighter skin linked to low-Communion in some cases.
- Intersectional analysis highlights darker-skinned males associated with the adjective 'poor' across models; darker females less represented.
- Ambiguity handling (AIAO) reduces explicit demographic cues in some cases but is not the main analysis focus.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.