Skip to main content
QUICK REVIEW

[논문 리뷰] Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation

Narek Tumanyan, Michal Geyer|arXiv (Cornell University)|2022. 11. 22.
Generative Adversarial Networks and Image Synthesis인용 수 13
한 줄 요약

이 논문은 미세조정 없이 사전 훈련된 텍스트-이미지 확산 모델을 활용하는 플러그 앤 플레이 프레임워크를 소개한다. 지도 이미지에서 공간적 특징과 자기주의 맵을 확산 과정에 통합함으로써, 복잡한 텍스트 프롬프트에 정확히 대응하면서도 높은 정밀도의 구조 보존을 달성한다. 이는 기존의 P2P와 같은 방법보다 레이아웃 일관성과 시각적 품질에서 뛰어나다.

ABSTRACT

Large-scale text-to-image generative models have been a revolutionary breakthrough in the evolution of generative AI, allowing us to synthesize diverse images that convey highly complex visual concepts. However, a pivotal challenge in leveraging such models for real-world content creation tasks is providing users with control over the generated content. In this paper, we present a new framework that takes text-to-image synthesis to the realm of image-to-image translation -- given a guidance image and a target text prompt, our method harnesses the power of a pre-trained text-to-image diffusion model to generate a new image that complies with the target text, while preserving the semantic layout of the source image. Specifically, we observe and empirically demonstrate that fine-grained control over the generated structure can be achieved by manipulating spatial features and their self-attention inside the model. This results in a simple and effective approach, where features extracted from the guidance image are directly injected into the generation process of the target image, requiring no training or fine-tuning and applicable for both real or generated guidance images. We demonstrate high-quality results on versatile text-guided image translation tasks, including translating sketches, rough drawings and animations into realistic images, changing of the class and appearance of objects in a given image, and modifications of global qualities such as lighting and color.

연구 동기 및 목표

  • 모델 재학습 또는 미세조정 없이 고정밀도의 텍스트 유도 이미지-이미지 번역을 가능하게 하기 위해.
  • 비텍스트 지침 입력에서 이미지를 생성할 때 텍스트-이미지 확산 모델의 구조적 제어 부족 문제를 해결하기 위해.
  • 확산 모델 내부의 공간적 특징과 자기주의가 어떻게 구조적 정보를 인코딩하는지 탐구하기 위해.
  • 실제 이미지, 스케치, 애니메이션과 같은 다양한 지침 입력에 적용 가능한 통합 프레임워크를 개발하기 위해.
  • 지시 이미지의 레이아웃을 보존하는 것과 목표 텍스트 프롬프트에 충실하는 것 사이의 더 나은 균형을 이루기 위해.

제안 방법

  • DDIM 역전파를 사용하여 지도 이미지에서 공간적 특징과 자기주의 맵을 추출한다.
  • 생성 과정 중 사전 훈련된 텍스트-이미지 모델의 잠재 확산 과정에 이러한 특징을 직접 주입한다.
  • 지도 이미지의 특징 의미론과 주의 패턴을 유지함으로써 공간적 구조를 보존한다.
  • 사전 훈련된 모델의 내부 표현만을 사용한다—추가 학습이나 파rameter 업데이트가 필요하지 않다.
  • 확산 모델의 잠재 공간에서 작동하며, 여러 노이즈 제거 단계에서 특징을 수정하여 생성을 유도한다.
  • 모델의 내부 주의 메커니즘을 활용하여 번역 중에 세밀한 구조적 연관성을 유지한다.
Figure 2 : Plug-and-play Diffusion Features. (a) Our framework takes as input a guidance image and a text prompt describing the desired translation; the guidance image is inverted to initial noise ${\boldsymbol{x}}^{G}_{T}$ , which is then progressively denoised using DDIM sampling. During this proc
Figure 2 : Plug-and-play Diffusion Features. (a) Our framework takes as input a guidance image and a text prompt describing the desired translation; the guidance image is inverted to initial noise ${\boldsymbol{x}}^{G}_{T}$ , which is then progressively denoised using DDIM sampling. During this proc

실험 결과

연구 질문

  • RQ1사전 훈련된 텍스트-이미지 확산 모델의 중간 공간적 특징 내에서 구조적 및 의미적 레이아웃 정보는 어떻게 인코딩되는가?
  • RQ2지도 이미지의 공간적 특징과 자기주의 맵을 사용하여 미세조정 없이 생성된 이미지의 구조를 제어할 수 있는가?
  • RQ3특징 주입은 P2P와 같은 교차주의 조작과 비교해 레이아웃 정밀도 유지에 어떻게 영향을 미치는가?
  • RQ4공간적 특징과 자기주의를 동시에 주입했을 때의 영향은 구조 보존과 시각적 품질에 어떤가?
  • RQ5어떤 상황에서 이 방법이 실패하며, 특징 공간 내 의미적 대응에 의존할 경우의 한계는 무엇인가?

주요 결과

  • 정량적 평가에서 자기 유사성 거리가 낮아 보여, P2P보다 다중 편집 시나리오에서 특히 뛰어난 구조 보존 성능을 달성한다.
  • 공간적 특징과 자기주의를 동시에 주입하는 것이 고정밀도 레이아웃 전송에 필수적이다; 하나의 구성 요소를 제거하면 구조 정밀도가 크게 악화된다.
  • 특히 객체 카테고리나 스타일 변화와 같은 구조적 변화를 다룰 때, 텍스트2라이브, 디퓨전클립, 플렉시티와 같은 베이스라인보다 정량적·정성적 지표에서 모두 뛰어나다.
  • 복잡한 편집, 예를 들어 스케치를 사실적인 사진으로 변환하거나 객체 정체성과 외관을 변경하는 경우에도 높은 시각적 품질과 의미 일관성을 유지한다.
  • DDIM 역전파를 통해 효과적인 지도 이미지 인코딩이 가능하지만, 무늬가 없는 이미지에서는 종종 저주파 수준의 외관 아티팩트를 포착할 수 있다.
  • 지도 이미지와 목표 텍스트 간에 의미적 대응이 없을 경우, 예를 들어 임의의 색상 세그멘테이션 마스크일 경우 실패한다.
Figure 3 : Visualising diffusion features. We used a collection of 20 humanoid images (real and generated), and extracted spatial features from different decoder layers, at roughly 50% of the generation process ( $t=540$ ). For each block, we applied PCA on the extracted features across all images a
Figure 3 : Visualising diffusion features. We used a collection of 20 humanoid images (real and generated), and extracted spatial features from different decoder layers, at roughly 50% of the generation process ( $t=540$ ). For each block, we applied PCA on the extracted features across all images a

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.