Skip to main content
QUICK REVIEW

[논문 리뷰] ControlCom: Controllable Image Composition using Diffusion Model

Bo Zhang, Yuxuan Duan|arXiv (Cornell University)|2023. 08. 19.
Generative Adversarial Networks and Image Synthesis인용 수 5
한 줄 요약

ControlCom은 조명과 자세를 제어할 수 있는 2차원 지시 벡터를 사용하여 이미지 블렌딩, 조화, 시각 합성 및 생성적 구성 등을 통합하는 유일한 확산 모델을 제안한다. 전역 및 국소 임bedding 융합 전략을 두 단계로 적용함으로써 전경의 정밀도와 제어 가능성을 향상시켜 공개 및 실세계 벤치마크에서 기존 방법들을 능가한다.

ABSTRACT

Image composition targets at synthesizing a realistic composite image from a pair of foreground and background images. Recently, generative composition methods are built on large pretrained diffusion models to generate composite images, considering their great potential in image generation. However, they suffer from lack of controllability on foreground attributes and poor preservation of foreground identity. To address these challenges, we propose a controllable image composition method that unifies four tasks in one diffusion model: image blending, image harmonization, view synthesis, and generative composition. Meanwhile, we design a self-supervised training framework coupled with a tailored pipeline of training data preparation. Moreover, we propose a local enhancement module to enhance the foreground details in the diffusion model, improving the foreground fidelity of composite images. The proposed method is evaluated on both public benchmark and real-world data, which demonstrates that our method can generate more faithful and controllable composite images than existing approaches. The code and model will be available at https://github.com/bcmi/ControlCom-Image-Composition.

연구 동기 및 목표

  • 기존 확산 기반 이미지 구성 방법에서 전경 속성(조명 및 자세)을 선택적으로 조정할 수 있는 제어 기능의 부족을 해결하기 위해.
  • 확산 기반 생성에서 자주 손실되는 질감 및 외형 세부 정보를 향상시켜 전경 정체성 유지 성능을 개선하기 위해.
  • 블렌딩, 조화, 시각 합성 및 생성적 구성의 네 가지 관련된 이미지 구성 작업을 하나의 조건부 확산 모델로 통합하기 위해.
  • 다중 구성 작업의 공동 훈련을 위한 맞춤형 데이터 준비 파이pline과 함께 자체 학습 훈련 프레임워크를 개발하기 위해.
  • 전역 일致성과 국소 세부 정보 정밀도를 향상시키기 위한 이중 단계 특징 융합 메커니즘을 설계하기 위해.

제안 방법

  • 조명 또는 자세 조정이 필요한지 여부를 지정하기 위해 조건부 제어 입력으로 2차원 지시 벡터를 도입하여 네 가지 구성 작업을 통합 처리할 수 있도록 한다.
  • 단일 통합 아키텍처를 사용하여 이미지 블렌딩, 조화, 시각 합성 및 생성적 구성 작업을 공동으로 훈련할 수 있도록 자체 학습 훈련 프레임워크를 설계한다.
  • 이중 단계 특징 융합 전략을 제안한다: 먼저 전역 전경 임베딩을 융합하여 거친 복합 이미지를 생성하고, 이후 국소 임베딩을 융합하여 질감 및 외형 세부 정보를 정밀하게 보완한다.
  • 국소 임베딩에서 유도된 정렬된 전경 임베딩 맵을 생성하고, 이를 확산 모델의 중간 특징을 조절하는 데 사용하여 국소 세부 정보 정밀도를 향상시킨다.
  • 사전 훈련된 이미지 인코더를 통해 배경 및 전경 이미지 양쪽에 조건을 줌으로써 확산 프로세스를 지시 벡터 및 융합된 임베딩으로 안내한다.
  • 모델 훈련을 위해 네 가지 작업 전반에 걸쳐 실제적인 구성 시나리오를 시뮬레이션하는 맞춤형 파이pline을 통해 훈련 데이터를 준비함으로써 종단 간 자체 학습 학습을 가능하게 한다.
Figure 1 : Overview of our controllable image composition method. We unify four tasks in one diffusion model and enable control over the illumination and pose of the synthesized foreground objects with a 2-dim indicator vector.
Figure 1 : Overview of our controllable image composition method. We unify four tasks in one diffusion model and enable control over the illumination and pose of the synthesized foreground objects with a 2-dim indicator vector.

실험 결과

연구 질문

  • RQ1공통 아키텍처 하에 하나의 확산 모델이 이미지 블렌딩, 조화, 시각 합성 및 생성적 구성 작업을 효과적으로 통합할 수 있는가?
  • RQ22차원 제어 벡터는 정체성 또는 현실감을 떨어뜨리지 않고 전경 속성(조명 및 자세)을 선택적으로 조작할 수 있는가?
  • RQ3전역 융합 후 국소 융합을 거치는 이중 단계 융합 전략은 동시에 또는 단일 단계 융합보다 전경 세부 정보 정밀도를 향상시키는가?
  • RQ4전경 정밀도, 배경 유지 및 전체 이미지 품질 측면에서 제안된 방법은 최신 기준 방법들과 비교해 어떻게 성능을 내는가?
  • RQ5모델은 정제된 벤치마크를 넘어서 실제 세계의 구성 시나리오에도 일반화 가능한가?

주요 결과

  • COCOEE 벤치마크에서 ControlCom의 구성 버전은 최고의 FID(3.19)와 QS(77.84)를 기록하여 PbE 및 ObjectStitch를 능가하는 종합적 품질을 확보했다.
  • 블렌딩 버전은 가장 높은 CLIP_{fg} 점수(90.63)를 기록하여 전경 정체성 유지 능력이 뛰어났으며, 배경 일관성도 우수했다.
  • 조화 및 시각 합성 버전은 PbE 및 ObjectStitch와 비교해 유사하거나 더 낮은 LPIPS 및 LSSIM 점수를 기록하여 시각적 조화와 현실감 향상을 입증했다.
  • 정성적 결과는 ControlCom이 텍스트 기반 또는 기준 확산 모델보다 전경 질감 및 외형 세부 정보를 훨씬 더 잘 유지함을 보여주었다.
  • 사용자 연구 및 시각적 비교 결과, ControlCom은 조명과 자세에 대해 정밀하고 제어 가능한 편집을 가능하게 하여 더 자연스럽고 현실적인 복합 이미지를 생성함을 확인했다.
  • 새로 구성된 FOSCom 데이터셋을 통해 실세계 데이터에 대한 일반화 능력이 뛰어나다는 것이 검증되었으며, 합성 벤치마크를 넘어서도 강력한 성능을 보였다.
Figure 2 : Illustration of our ControlCom. Our model consists of two main components: a foreground encoder (a) that extracts hierarchical embeddings from foreground image, and a controllable generator (b) that synthesizes composite image with control over foreground illumination and pose using indic
Figure 2 : Illustration of our ControlCom. Our model consists of two main components: a foreground encoder (a) that extracts hierarchical embeddings from foreground image, and a controllable generator (b) that synthesizes composite image with control over foreground illumination and pose using indic

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.