[论文解读] Design Guidelines for Prompt Engineering Text-to-Image Generative Models
本论文分析提示词措辞、随机种子、迭代长度,以及风格/主题选择如何影响文本到图像生成(超过5个实验与5493次生成),并推导出更好结果的实际设计准则。
Text-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of generations, they also must engage in brute-force trial and error with the text prompt when the result quality is poor. We conduct a study exploring what prompt keywords and model hyperparameters can help produce coherent outputs. In particular, we study prompts structured to include subject and style keywords and investigate success and failure modes of these prompts. Our evaluation of 5493 generations over the course of five experiments spans 51 abstract and concrete subjects as well as 51 abstract and figurative styles. From this evaluation, we present design guidelines that can help people produce better outcomes from text-to-image generative models.
研究动机与目标
- 研究提示词关键词和模型超参数如何影响文本到图像生成的质量与连贯性。
- 系统性评估以“SUBJECT in the style of STYLE”结构的提示词在大量主体与风格上的表现。
- 识别成功与失败模式,并将研究发现转化为可供最终用户执行的可操作设计准则。
提出的方法
- 在实验1中,使用 VQGAN+CLIP,256x256 图像,每张图像进行 300 个优化步骤,在 51 个主体和 12 种风格上生成提示。
- 对每个主体-风格对测试九种提示词表达形式的排列,以评估提示措辞的影响。
- 扰动种子(随机初始化),并分析不同种子是否产生显著不同的生成结果。
- 改变优化长度(迭代次数)以确定其与感知质量的相关性。
- 在 12 个主体上测试 51 种风格,以评估风格表示的广度及潜在偏差。
- 使用人工评注对生成结果打分,并进行统计检验(Fisher’s exact test、Chi-square、Cohen’s kappa)以确定显著性。

实验结果
研究问题
- RQ1对同一关键词的不同措辞是否会产生显著不同的生成结果?
- RQ2在固定提示下,随机种子是否显著影响生成质量?
- RQ3优化长度如何影响生成质量与用户偏好?
- RQ4模型对广泛风格的表示能力有多强,是否存在风格偏差?
- RQ5主体和风格如何交互影响生成结果?
主要发现
- 提示排列:在九种提示变体之间没有显著差异;应聚焦于主体/风格关键词,而非连接词。
- 种子变化:种子选择显著影响生成质量;建议每个提示生成 3–9 个种子以捕捉变异性。
- 优化长度:较短的运行(100–500 次迭代)常被偏好;建议 300 次迭代作为一个良好默认值。
- 风格广度:模型在 51 种风格上的表现不同,可识别的成功模式包括颜色、技法、空间关系和主题符号;观察到特定风格的偏差。
- 总体而言,风格和主体与模型能力互动,能够实现定性的成功模式,但在不同风格间存在差异。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。