[Paper Review] Holistic Evaluation of Text-To-Image Models
The paper introduces HEIM, a holistic benchmark evaluating 26 text-to-image models across 12 aspects using 62 scenarios and human+automated metrics, revealing diverse strengths across models and generally weak correlations between human and automated measures.
The stunning qualitative improvement of recent text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding of their capabilities and risks. To fill this gap, we introduce a new benchmark, Holistic Evaluation of Text-to-Image Models (HEIM). Whereas previous evaluations focus mostly on text-image alignment and image quality, we identify 12 aspects, including text-image alignment, image quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency. We curate 62 scenarios encompassing these aspects and evaluate 26 state-of-the-art text-to-image models on this benchmark. Our results reveal that no single model excels in all aspects, with different models demonstrating different strengths. We release the generated images and human evaluation results for full transparency at https://crfm.stanford.edu/heim/v1.1.0 and the code at https://github.com/stanford-crfm/helm, which is integrated with the HELM codebase.
Motivation & Objective
- Establish a holistic benchmark to evaluate text-to-image models beyond image quality and alignment.
- Assess 12 dimensions including bias, toxicity, fairness, multilinguality, and efficiency for practical deployment.
- Provide standardized, transparent evaluation across a broad set of models and scenarios.
- Offer insights into how different models perform across diverse, realistic prompting tasks.
Proposed method
- Define 12 evaluation aspects and 62 prompting scenarios (prompts and references) to cover broad capabilities and risks.
- Use 25 metrics, combining human judgments (crowdsourced) and automated measures, to assess each aspect.
- Standardize evaluation across 26 recent text-to-image models under zero-shot prompting and common adaptation strategies.
- Curate new scenarios and metrics for underexplored aspects like originality, aesthetics, bias, fairness, multilinguality, robustness, and efficiency.
- Release generated images, human results, and evaluation code for transparency and reproducibility.

Experimental results
Research questions
- RQ1How do state-of-the-art text-to-image models perform across a broad set of holistic aspects beyond traditional alignment and quality?
- RQ2What are the relationships between human judgments and automated metrics in evaluating these models?
- RQ3Which models exhibit strengths or weaknesses across different aspects, and what ethical/societal risks emerge from current capabilities?
- RQ4How do factors like multilinguality, robustness, and efficiency influence practical deployment of text-to-image models.
Key findings
- No single model excels in all aspects; different models have distinct strengths (e.g., DALL-E 2 for alignment, Openjourney for aesthetics, minDALL-E and Safe Stable Diffusion for bias/toxicity mitigation).
- Correlations between human and automated metrics are generally weak, especially for photorealism and aesthetics, underscoring the value of human evaluation.
- Several aspects require more attention: reasoning and multilinguality lag behind others, while originality, toxicity, and bias raise ethical/legal concerns.
- Prompt engineering improves visual appeal, with Promptist+Stable Diffusion outperforming in aesthetics while maintaining alignment.
- Art-fine-tuned models excel in aesthetics or realism but may trade off other aspects like alignment or bias mitigation.
- DALL-E 2 often leads in human-aligned performance, but no model dominates all societal or multilingual aspects; different models offer complementary strengths.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.