Skip to main content
QUICK REVIEW

[논문 리뷰] Design Guidelines for Prompt Engineering Text-to-Image Generative Models

Vivian Liu, Lydia B. Chilton|arXiv (Cornell University)|2021. 09. 14.
Aesthetic Perception and Analysis인용 수 32
한 줄 요약

이 논문은 프롬프트 문구, 난수 시드, 반복 길이, 스타일/주제 선택이 텍스트-이미지 생성에 어떤 영향을 미치는지 분석하고(5개 실험 및 5493개의 생성에 대해) 더 나은 결과를 위한 실용적 설계 지침을 도출한다.

ABSTRACT

Text-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of generations, they also must engage in brute-force trial and error with the text prompt when the result quality is poor. We conduct a study exploring what prompt keywords and model hyperparameters can help produce coherent outputs. In particular, we study prompts structured to include subject and style keywords and investigate success and failure modes of these prompts. Our evaluation of 5493 generations over the course of five experiments spans 51 abstract and concrete subjects as well as 51 abstract and figurative styles. From this evaluation, we present design guidelines that can help people produce better outcomes from text-to-image generative models.

연구 동기 및 목표

  • 프롬프트 키워드와 모델 하이퍼파라미터가 텍스트-투-이미지 생성의 품질과 일관성에 어떤 영향을 미치는지 조사한다.
  • 여러 주제와 스타일에 걸쳐 'SUBJECT in the style of STYLE'로 구성된 프롬 prompts를 체계적으로 평가한다.
  • 성공 모드와 실패 모드를 식별하고 연구 결과를 최종 사용자를 위한 실행 가능한 설계 지침으로 도출한다.

제안 방법

  • Experiment 1에서 51개의 주제와 12개의 스타일에 걸쳐 프롬프트를 생성하기 위해 256x256 이미지의 VQGAN+CLIP을 사용하고, 이미지당 300개의 최적화 단계를 수행한다.
  • 각 주제-스타일 쌍에 대해 프롬프트 구성의 9가지 배열을 테스트하여 프롬프트 문구의 영향을 평가한다.
  • 시드(랜덤 초기화)를 섞어 서로 다른 시드가 유의하게 다른 생성을 만들어내는지 분석한다.
  • 최적화 길이(반복 횟수)를 변화시켜 인지된 품질과의 상관관계를 결정한다.
  • 12개 주제에 걸쳐 51개 스타일을 테스트하여 스타일 표현의 폭과 잠재적 편향을 평가한다.
  • 인간 평가자에 의해 생성물을 주석하고 통계적 검정(Fisher’s exact test, Chi-square, Cohen’s kappa)을 수행하여 유의성을 판단한다.
Figure 1. An example grid of text-to-image generations generated from the following prompt template: ”SUBJECT in the style of STYLE”. We analyze over 5000 generations in a series of five experiments involving 51 subjects and 51 styles to study what prompt parameters and hyperparameters can help peop
Figure 1. An example grid of text-to-image generations generated from the following prompt template: ”SUBJECT in the style of STYLE”. We analyze over 5000 generations in a series of five experiments involving 51 subjects and 51 styles to study what prompt parameters and hyperparameters can help peop

실험 결과

연구 질문

  • RQ1동일한 키워드의 다른 표현이 유의하게 다른 생성을 야기하는가?
  • RQ2고정된 프롬프트에서 난수 시드가 생성 품질에 유의하게 영향을 미치는가?
  • RQ3최적화 길이가 생성 품질과 사용자 선호도에 어떤 영향을 미치는가?
  • RQ4모델이 광범위한 스타일을 얼마나 잘 표현할 수 있으며 스타일 편향이 있는가?
  • RQ5주제와 스타일이 어떻게 상호작용하여 생성 결과에 영향을 미치는가?

주요 결과

  • 프롬프트 순열: 9가지 프롬프트 변형 간 유의한 차이가 없으며, 연결어보다는 주제/스타일 키워드에 집중한다.
  • 시드 변화: 시드 선택이 생성 품질에 유의하게 영향을 미치며, 변동성을 포착하기 위해 프롬프트당 3–9개의 시드를 생성하는 것을 권장한다.
  • 최적화 길이: 짧은 실행(100–500 반복)이 종종 선호되며, 300 반복이 좋은 기본값으로 제안된다.
  • 스타일의 폭: 모델 성능은 51개 스타일에 걸쳐 다르며, 색상, 기법, 공간 관계, 모티프에서 식별 가능한 성공 모드가 있고, 스타일별 편향이 관찰된다.
  • 전반적으로 스타일과 주제는 모델의 역량과 상호 작용하여 질적 성공 모드를 가능하게 하지만 스타일에 따라 변동이 있다.
Figure 2. For Experiment 1, annotators judged 3x3 grids where generations from different prompt permutations were arranged randomly. Annotators evaluated 143 grids of generations for significantly better generations as well significantly worse generations (outliers in generation quality). We found n
Figure 2. For Experiment 1, annotators judged 3x3 grids where generations from different prompt permutations were arranged randomly. Annotators evaluated 143 grids of generations for significantly better generations as well significantly worse generations (outliers in generation quality). We found n

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.