Skip to main content
QUICK REVIEW

[Paper Review] Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation

Yixin Wan, Arjun Subramonian|arXiv (Cornell University)|Apr 1, 2024
Data Visualization and Analytics4 citations
TL;DR

This survey provides the first comprehensive analysis of bias in text-to-image (T2I) generation models, examining definitions, evaluation methods, and mitigation strategies across three dimensions: gender presentation, skintone, and geo-cultural bias. It reveals significant gaps—especially in under-explored geo-cultural bias, non-binary representation, and lack of unified evaluation frameworks—while advocating for human-centric, adaptive, and lifelong mitigation approaches to build fairer AI systems.

ABSTRACT

The recent advancement of large and powerful models with Text-to-Image (T2I) generation abilities -- such as OpenAI's DALLE-3 and Google's Gemini -- enables users to generate high-quality images from textual prompts. However, it has become increasingly evident that even simple prompts could cause T2I models to exhibit conspicuous social bias in generated images. Such bias might lead to both allocational and representational harms in society, further marginalizing minority groups. Noting this problem, a large body of recent works has been dedicated to investigating different dimensions of bias in T2I systems. However, an extensive review of these studies is lacking, hindering a systematic understanding of current progress and research gaps. We present the first extensive survey on bias in T2I generative models. In this survey, we review prior studies on dimensions of bias: Gender, Skintone, and Geo-Culture. Specifically, we discuss how these works define, evaluate, and mitigate different aspects of bias. We found that: (1) while gender and skintone biases are widely studied, geo-cultural bias remains under-explored; (2) most works on gender and skintone bias investigated occupational association, while other aspects are less frequently studied; (3) almost all gender bias works overlook non-binary identities in their studies; (4) evaluation datasets and metrics are scattered, with no unified framework for measuring biases; and (5) current mitigation methods fail to resolve biases comprehensively. Based on current limitations, we point out future research directions that contribute to human-centric definitions, evaluations, and mitigation of biases. We hope to highlight the importance of studying biases in T2I systems, as well as encourage future efforts to holistically understand and tackle biases, building fair and trustworthy T2I technologies for everyone.

Motivation & Objective

  • To systematically review and synthesize existing research on bias in text-to-image (T2I) generation models across three core dimensions: gender presentation, skintone, and geo-cultural bias.
  • To identify and analyze how prior works define, evaluate, and mitigate bias in T2I systems, highlighting inconsistencies and methodological shortcomings.
  • To expose critical research gaps, including the underrepresentation of geo-cultural bias, lack of attention to non-binary identities, and absence of unified evaluation frameworks.
  • To advocate for future research directions centered on human-centric definitions, dynamic evaluation, and adaptive, lifelong mitigation strategies that evolve with societal norms.
  • To support the development of fairer, more trustworthy T2I technologies by guiding researchers and policymakers toward ethically responsible AI design and deployment.

Proposed method

  • Conducted a systematic literature review of 36 peer-reviewed papers on bias in T2I models, focusing on gender presentation, skintone, and geo-cultural bias.
  • Categorized and analyzed studies based on their conceptualization of bias, evaluation methodologies, and mitigation techniques across the three bias dimensions.
  • Mapped definitions of bias to specific aspects such as occupational association, power dynamics, stereotypical objects, and image quality to identify thematic clusters.
  • Evaluated the diversity and consistency of evaluation metrics and datasets used across studies, identifying fragmentation and lack of standardization.
  • Proposed a framework for future research emphasizing adaptive, lifelong mitigation strategies that evolve with changing societal norms and community feedback.
  • Highlighted ethical risks in identity classification, including reliance on visual proxies for gender, race, and culture, and cautioned against misuse in surveillance or oppressive applications.

Experimental results

Research questions

  • RQ1How do existing studies define bias in text-to-image generation models across gender presentation, skintone, and geo-cultural dimensions?
  • RQ2What evaluation metrics and datasets are commonly used to measure bias in T2I models, and to what extent is there standardization across studies?
  • RQ3What mitigation strategies have been proposed, and how effective are they in addressing multiple dimensions of bias in a holistic manner?
  • RQ4Why is geo-cultural bias significantly under-explored compared to gender and skintone bias in current research?
  • RQ5How can future research develop human-centric, adaptive, and lifelong approaches to bias mitigation that respond to evolving societal values and community needs?

Key findings

  • Gender and skintone biases are widely studied, particularly in the context of occupational associations, while geo-cultural bias remains significantly under-explored in the literature.
  • Most gender bias research fails to account for non-binary gender identities, resulting in exclusionary representations in T2I model outputs.
  • There is no unified evaluation framework for bias in T2I models, with metrics and datasets varying widely across studies, leading to inconsistent and non-comparable results.
  • Current mitigation methods are largely ineffective at resolving bias holistically, often failing to address multiple dimensions of bias simultaneously or in a sustainable way.
  • Evaluation practices are compromised by biased classification methods—both human and automated—leading to propagation of stereotypes in bias measurement itself.
  • The lack of dynamic, adaptive, and lifelong mitigation strategies limits the long-term fairness of T2I systems, especially as societal norms and cultural representations evolve over time.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.