[论文解读] Holistic Evaluation of Text-To-Image Models
本文提出 HEIM,一种全面基准,使用 62 个场景和人类+自动化评测,在 12 个方面评估 26 种文本到图像模型,揭示各模型的多样化优势,以及人类与自动化评测之间通常薄弱的相关性。
The stunning qualitative improvement of recent text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding of their capabilities and risks. To fill this gap, we introduce a new benchmark, Holistic Evaluation of Text-to-Image Models (HEIM). Whereas previous evaluations focus mostly on text-image alignment and image quality, we identify 12 aspects, including text-image alignment, image quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency. We curate 62 scenarios encompassing these aspects and evaluate 26 state-of-the-art text-to-image models on this benchmark. Our results reveal that no single model excels in all aspects, with different models demonstrating different strengths. We release the generated images and human evaluation results for full transparency at https://crfm.stanford.edu/heim/v1.1.0 and the code at https://github.com/stanford-crfm/helm, which is integrated with the HELM codebase.
研究动机与目标
- 建立一个全面基准,用于评估文本到图像模型,超越图像质量和对齐度。
- 评估包括偏见、毒性、公平性、多语言性和效率在内的 12 个维度,便于实际部署。
- 在广泛的模型和场景上提供标准化、透明的评估。
- 提供不同模型在多样、真实的提示任务中的表现洞见。
提出的方法
- 定义 12 个评估方面和 62 个提示场景(提示和参考),覆盖广泛能力与风险。
- 使用 25 个指标,结合人类判断(众包)与自动化度量来评估每个方面。
- 在 26 个最近的文本到图像模型上,采用零-shot 提示和常见适配策略,标准化评估。
- 为未充分探索的方面如原创性、审美、偏见、公平性、多语言性、鲁棒性和效率,策划新场景和指标。
- 发布生成的图像、人类结果和评测代码,以实现透明性和可重复性。

实验结果
研究问题
- RQ1最先进的文本到图像模型在超越传统对齐与质量的广泛综合方面表现如何?
- RQ2在评估这些模型时,人类判断与自动化指标之间的关系是什么?
- RQ3哪些模型在不同方面表现出优势或劣势,当前能力带来哪些伦理/社会风险?
- RQ4多语言性、鲁棒性和效率等因素如何影响文本到图像模型的实际部署?
主要发现
- 没有单一模型在所有方面都出类拔萃;不同模型具有各自的优势(例如:DALL-E 2 在对齐方面,Openjourney 在美学方面,minDALL-E 与 Safe Stable Diffusion 在偏见/毒性缓解方面)。
- 人类与自动化指标之间的相关性通常很弱,尤其在照片写实性和美学方面,这凸显了人类评估的价值。
- 有若干方面需要更多关注:推理能力和多语言性落后于其他方面,而原创性、毒性和偏见引发伦理/法律方面的担忧。
- 提示工程提升了视觉吸引力,其中 Promptist+Stable Diffusion 在美学方面表现优于其他方案,同时保持对齐。
- 艺术风格微调的模型在美学或真实感方面表现出色,但可能在对齐或偏见缓解等其他方面作出权衡。
- DALL-E 2 常在与人类对齐的表现方面领先,但没有模型在所有社会或多语言方面都占主导;不同模型提供互补的优势。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。