[Paper Review] Towards Equitable Representation in Text-to-Image Synthesis Models with the Cross-Cultural Understanding Benchmark (CCUB) Dataset
This paper proposes a culturally-aware text-to-image synthesis approach using the Cross-Cultural Understanding Benchmark (CCUB) dataset—1,095 curated image-text pairs from eight cultures—to reduce bias and offensiveness in Stable Diffusion. By fine-tuning the model on CCUB data and using GPT-3 to augment prompts with cultural context, the method improves cultural relevance and reduces offensiveness while maintaining image quality.
It has been shown that accurate representation in media improves the well-being of the people who consume it. By contrast, inaccurate representations can negatively affect viewers and lead to harmful perceptions of other cultures. To achieve inclusive representation in generated images, we propose a culturally-aware priming approach for text-to-image synthesis using a small but culturally curated dataset that we collected, known here as Cross-Cultural Understanding Benchmark (CCUB) Dataset, to fight the bias prevalent in giant datasets. Our proposed approach is comprised of two fine-tuning techniques: (1) Adding visual context via fine-tuning a pre-trained text-to-image synthesis model, Stable Diffusion, on the CCUB text-image pairs, and (2) Adding semantic context via automated prompt engineering using the fine-tuned large language model, GPT-3, trained on our CCUB culturally-aware text data. CCUB dataset is curated and our approach is evaluated by people who have a personal relationship with that particular culture. Our experiments indicate that priming using both text and image is effective in improving the cultural relevance and decreasing the offensiveness of generated images while maintaining quality.
Motivation & Objective
- Address the lack of equitable cultural representation in text-to-image synthesis models trained on biased, large-scale datasets.
- Mitigate harmful stereotypes and offensiveness in AI-generated images, especially for underrepresented cultural groups.
- Develop a scalable method to customize pre-trained models for culturally relevant image generation without retraining from scratch.
- Evaluate cultural relevance and offensiveness using native cultural experts to ensure authentic feedback.
- Create a publicly available benchmark dataset (CCUB) to support future research in culturally aware AI generation.
Proposed method
- Curated the Cross-Cultural Understanding Benchmark (CCUB) dataset with 1,095 image-text pairs across eight countries, collected by native cultural experts.
- Fine-tuned the Stable Diffusion model on CCUB data to learn culture-specific visual concepts and improve cultural alignment.
- Trained a GPT-3 model on CCUB’s culturally aware text data to automatically augment input prompts with relevant cultural details (e.g., local customs, tools, attire).
- Combined fine-tuned image generation with GPT-3 prompt augmentation to enhance both visual and semantic cultural relevance.
- Evaluated the system using a survey with 72 native participants from five countries, comparing image quality, cultural alignment, and offensiveness.
- Used a baseline of simply appending the country name to the prompt to isolate the impact of the proposed techniques.
Experimental results
Research questions
- RQ1Can fine-tuning a pre-trained text-to-image model on a small, culturally curated dataset improve cultural relevance in generated images?
- RQ2Does automated prompt augmentation using a fine-tuned LLM enhance cultural specificity without degrading image quality?
- RQ3How does combining visual fine-tuning and semantic prompt augmentation compare to baseline methods in reducing offensiveness and improving cultural alignment?
- RQ4Can native cultural experts reliably assess the cultural accuracy and offensiveness of AI-generated images?
- RQ5To what extent does nation-based cultural representation in the dataset generalize to real-world cultural diversity and intersectionality?
Key findings
- The combined approach of visual fine-tuning and prompt augmentation significantly outperformed the baseline in cultural alignment, with a 38% reduction in offensiveness compared to the baseline.
- Fine-tuning alone improved cultural alignment and reduced offensiveness, but had no significant effect on text-image alignment.
- Prompt augmentation alone increased cultural alignment but did not reduce offensiveness and slightly decreased text-image alignment due to mismatched prompt-content generation.
- The combined method achieved the best balance—improving cultural alignment while maintaining or slightly improving text-image alignment and reducing offensiveness.
- Native evaluators consistently rated images generated with the combined method as more authentic and less stereotypical than those from the baseline.
- The CCUB dataset demonstrated high cultural specificity and was effective in guiding models toward accurate, contextually appropriate representations of diverse cultures.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.