Skip to main content
QUICK REVIEW

[Paper Review] Holistic Evaluation of Text-To-Image Models

Tong Lee, Michihiro Yasunaga|arXiv (Cornell University)|Nov 7, 2023
Multimedia Communication and Technology13 citations
TL;DR

The paper introduces HEIM, a holistic benchmark evaluating 26 text-to-image models across 12 aspects using 62 scenarios and human+automated metrics, revealing diverse strengths across models and generally weak correlations between human and automated measures.

ABSTRACT

The stunning qualitative improvement of recent text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding of their capabilities and risks. To fill this gap, we introduce a new benchmark, Holistic Evaluation of Text-to-Image Models (HEIM). Whereas previous evaluations focus mostly on text-image alignment and image quality, we identify 12 aspects, including text-image alignment, image quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency. We curate 62 scenarios encompassing these aspects and evaluate 26 state-of-the-art text-to-image models on this benchmark. Our results reveal that no single model excels in all aspects, with different models demonstrating different strengths. We release the generated images and human evaluation results for full transparency at https://crfm.stanford.edu/heim/v1.1.0 and the code at https://github.com/stanford-crfm/helm, which is integrated with the HELM codebase.

Motivation & Objective

  • Establish a holistic benchmark to evaluate text-to-image models beyond image quality and alignment.
  • Assess 12 dimensions including bias, toxicity, fairness, multilinguality, and efficiency for practical deployment.
  • Provide standardized, transparent evaluation across a broad set of models and scenarios.
  • Offer insights into how different models perform across diverse, realistic prompting tasks.

Proposed method

  • Define 12 evaluation aspects and 62 prompting scenarios (prompts and references) to cover broad capabilities and risks.
  • Use 25 metrics, combining human judgments (crowdsourced) and automated measures, to assess each aspect.
  • Standardize evaluation across 26 recent text-to-image models under zero-shot prompting and common adaptation strategies.
  • Curate new scenarios and metrics for underexplored aspects like originality, aesthetics, bias, fairness, multilinguality, robustness, and efficiency.
  • Release generated images, human results, and evaluation code for transparency and reproducibility.
Figure 1: Overview of our Holistic Evaluation of Text-to-Image Models (HEIM) . While existing benchmarks focus on limited aspects such as image quality and alignment with text, rely on automated metrics that may not accurately reflect human judgment, and evaluate limited models, HEIM takes a holisti
Figure 1: Overview of our Holistic Evaluation of Text-to-Image Models (HEIM) . While existing benchmarks focus on limited aspects such as image quality and alignment with text, rely on automated metrics that may not accurately reflect human judgment, and evaluate limited models, HEIM takes a holisti

Experimental results

Research questions

  • RQ1How do state-of-the-art text-to-image models perform across a broad set of holistic aspects beyond traditional alignment and quality?
  • RQ2What are the relationships between human judgments and automated metrics in evaluating these models?
  • RQ3Which models exhibit strengths or weaknesses across different aspects, and what ethical/societal risks emerge from current capabilities?
  • RQ4How do factors like multilinguality, robustness, and efficiency influence practical deployment of text-to-image models.

Key findings

  • No single model excels in all aspects; different models have distinct strengths (e.g., DALL-E 2 for alignment, Openjourney for aesthetics, minDALL-E and Safe Stable Diffusion for bias/toxicity mitigation).
  • Correlations between human and automated metrics are generally weak, especially for photorealism and aesthetics, underscoring the value of human evaluation.
  • Several aspects require more attention: reasoning and multilinguality lag behind others, while originality, toxicity, and bias raise ethical/legal concerns.
  • Prompt engineering improves visual appeal, with Promptist+Stable Diffusion outperforming in aesthetics while maintaining alignment.
  • Art-fine-tuned models excel in aesthetics or realism but may trade off other aspects like alignment or bias mitigation.
  • DALL-E 2 often leads in human-aligned performance, but no model dominates all societal or multilingual aspects; different models offer complementary strengths.
Figure 2: Standardized evaluation . Prior to HEIM ( top panel ), the evaluation of image generation models was not comprehensive: six of our 12 core aspects were not evaluated on existing models, and only 11% of the total evaluation space was studied (the percentage of ✓in the matrix of aspects $\ti
Figure 2: Standardized evaluation . Prior to HEIM ( top panel ), the evaluation of image generation models was not comprehensive: six of our 12 core aspects were not evaluated on existing models, and only 11% of the total evaluation space was studied (the percentage of ✓in the matrix of aspects $\ti

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.