[Paper Review] Imagen 3
Imagen 3 is a state-of-the-art latent diffusion model that generates high-resolution (1024×1024) text-to-image outputs with exceptional photorealism and prompt adherence. Trained on a filtered, diverse dataset with synthetic captions from Gemini models, it outperforms SOTA models like DALL·E 3 and Stable Diffusion 3 in human evaluations, achieving the highest Elo scores in overall preference, visual quality, and prompt-image alignment.
We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of evaluation. In addition, we discuss issues around safety and representation, as well as methods we used to minimize the potential harm of our models.
Motivation & Objective
- To develop a text-to-image diffusion model that generates high-fidelity, detailed images from complex natural language prompts.
- To improve alignment between long and intricate prompts and generated image content, especially in compositional and stylistic details.
- To ensure safety and responsible generation by filtering harmful data and minimizing bias and overfitting in training.
- To establish a new benchmark in text-to-image generation through rigorous human and automatic evaluation against SOTA models.
- To minimize potential harms through multi-stage data filtering, deduplication, and responsible data sourcing practices.
Proposed method
- Training a latent diffusion model on a large-scale, multi-stage filtered dataset comprising real images, original captions, and synthetic captions from Gemini models.
- Applying a multi-stage data filtering pipeline to remove unsafe, low-quality, and AI-generated images, and to reduce data redundancy via deduplication and down-weighting.
- Generating synthetic captions using multiple Gemini models with diverse instructions to enhance linguistic diversity and quality.
- Using side-by-side human evaluations with Elo scoring to compare model outputs across five quality aspects: overall preference, prompt-image alignment, visual appeal, detailed alignment, and numerical reasoning.
- Employing external benchmarks such as GenAI-Bench, DOCCI-Test-Pivots, and GeckoNum to evaluate generalization and reasoning capabilities.
- Conducting evaluations via public APIs for external models and using released images for DALL·E 3 to ensure consistency and fairness.
Experimental results
Research questions
- RQ1Can Imagen 3 generate images that are preferred over other SOTA models in human evaluations across multiple quality dimensions?
- RQ2To what extent does Imagen 3 maintain detailed prompt-image alignment, especially for complex, long-form descriptions?
- RQ3How does Imagen 3 perform in numerical reasoning tasks, such as accurately depicting specified numbers of objects in an image?
- RQ4What impact do multi-stage data filtering and synthetic captioning have on reducing bias and improving model generalization?
- RQ5How effective are the safety and responsibility measures in minimizing harmful or harmful-reinforcing outputs?
Key findings
- Imagen 3 achieved the highest Elo score in overall preference (1,115) on the GenAI-Bench benchmark, significantly outperforming DALL·E 3 (1,078) and Stable Diffusion 3.5 Large (1,059).
- In visual quality evaluation, Imagen 3-002 achieved an Elo score of 1,135, surpassing all other models, including DALL·E 3 (1,112) and Stable Diffusion 3.5 Large (1,104).
- For prompt-image alignment, Imagen 3-002 scored 1,106 on the GenAI-Bench set, exceeding DALL·E 3 (1,066) and Midjourney v6 (1,063).
- On the GeckoNum benchmark, Imagen 3 demonstrated strong numerical reasoning, correctly depicting object counts in complex scenes with high accuracy.
- The model showed superior performance in detailed alignment tasks using DOCCI-Test-Pivots, maintaining complex compositional elements from prompts.
- Human raters consistently preferred Imagen 3 across all evaluation sets, with over 366,000 ratings collected from 3,225 distinct raters to ensure reliability and reduce bias.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.