Skip to main content
QUICK REVIEW

[論文レビュー] Design Guidelines for Prompt Engineering Text-to-Image Generative Models

Vivian Liu, Lydia B. Chilton|arXiv (Cornell University)|Sep 14, 2021
Aesthetic Perception and Analysis被引用数 32
ひとこと要約

この論文は、プロンプトの表現、乱数種、反復長、スタイル/主題の選択がテキスト対画像生成にどのように影響するかを、5つの実験と5493世代に及ぶ分析を通じて検討し、より良い成果のための実践的な設計ガイドラインを導出します。

ABSTRACT

Text-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of generations, they also must engage in brute-force trial and error with the text prompt when the result quality is poor. We conduct a study exploring what prompt keywords and model hyperparameters can help produce coherent outputs. In particular, we study prompts structured to include subject and style keywords and investigate success and failure modes of these prompts. Our evaluation of 5493 generations over the course of five experiments spans 51 abstract and concrete subjects as well as 51 abstract and figurative styles. From this evaluation, we present design guidelines that can help people produce better outcomes from text-to-image generative models.

研究の動機と目的

  • プロンプトキーワードとモデルのハイパーパラメータが、テキストから画像生成の品質と一貫性にどのように影響するかを調査する。
  • 「SUBJECT in the style of STYLE」の形式で構成されたプロンプトを、多数の主題とスタイルに対して体系的に評価する。
  • 成功と失敗のモードを特定し、発見をエンドユーザー向けの実用的な設計ガイドラインへ翻訳する。

提案手法

  • Experiment 1では、VQGAN+CLIPを用い、256x256画像、1枚あたり300の最適化ステップで、51の主題と12のスタイルに対してプロンプトを生成する。
  • 各主題-スタイルペアについて、プロンプト表現の9つの置換をテストし、プロンプト表現の影響を評価する。
  • 乱数初期値(シード)を摂動させ、異なるシードが有意に異なる生成を生み出すかを分析する。
  • 最適化の長さ(反復回数)を変化させ、知覚品質との相関を把握する。
  • 12の主題にわたり51のスタイルをテストして、スタイル表現の広がりと潜在的なバイアスを評価する。
  • 人間の評価者で生成物を注釈付けし、有意性を判断するために統計検定(Fisher’s exact test、Chi-square、Cohen’s kappa)を実施する。
Figure 1. An example grid of text-to-image generations generated from the following prompt template: ”SUBJECT in the style of STYLE”. We analyze over 5000 generations in a series of five experiments involving 51 subjects and 51 styles to study what prompt parameters and hyperparameters can help peop
Figure 1. An example grid of text-to-image generations generated from the following prompt template: ”SUBJECT in the style of STYLE”. We analyze over 5000 generations in a series of five experiments involving 51 subjects and 51 styles to study what prompt parameters and hyperparameters can help peop

実験結果

リサーチクエスチョン

  • RQ1同じキーワードの異なる言い回しは、統計的に有意に異なる生成を生み出すか?
  • RQ2固定プロンプトで、乱数種は生成品質に有意に影響するか?
  • RQ3最適化の長さは生成品質とユーザー好みにどのように影響するか?
  • RQ4モデルは大きなスタイルの広がりをどれだけ表現できるか、そしてスタイルのバイアスはあるか?
  • RQ5主題とスタイルは生成結果にどのように相互作用して影響するか?

主な発見

  • プロンプトの置換: 9つのプロンプト変種間で有意な差はなし。接続語よりも主題/スタイルのキーワードに焦点を当てるべき。
  • シードの変動: シードの選択が生成品質に有意に影響する。変動を捉えるために、各プロンプトにつき3–9個のシードを生成することを推奨。
  • 最適化の長さ: 短い実行(100–500反復)が好まれることが多い。300反復を良いデフォルトとして推奨。
  • スタイルの広がり: 51のスタイルでモデルの性能は異なり、色、技法、空間関係、モチーフなどの成功モードが特定される。観察されたスタイル特有のバイアス。
  • 全体として、スタイルと主題はモデルの能力と相互作用し、定性的な成功モードを可能にするが、スタイルごとにばらつきがある。
Figure 2. For Experiment 1, annotators judged 3x3 grids where generations from different prompt permutations were arranged randomly. Annotators evaluated 143 grids of generations for significantly better generations as well significantly worse generations (outliers in generation quality). We found n
Figure 2. For Experiment 1, annotators judged 3x3 grids where generations from different prompt permutations were arranged randomly. Annotators evaluated 143 grids of generations for significantly better generations as well significantly worse generations (outliers in generation quality). We found n

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。