[논문 리뷰] Boomerang: Local sampling on image manifolds using diffusion models
보머랑은 사전 훈련된 확산 모델을 사용하여 입력 이미지에 제어 가능한 노이즈를 추가한 후 정방향 확산 과정을 뒤집음으로써 이미지 매니폴드에서 국소적 샘플링을 가능하게 한다. 이는 원본에 가까운 다양한 실재감 있는 변형을 생성하며, 모델 미세조정 없이 유사성과 확률적 다양성을 제어할 수 있다. 이는 개인 정보 보호 데이터 생성, 데이터 증강, 단일 GPU에서의 8배 초해상도 처리 등에 활용 가능하다.
The inference stage of diffusion models can be seen as running a reverse-time diffusion stochastic differential equation, where samples from a Gaussian latent distribution are transformed into samples from a target distribution that usually reside on a low-dimensional manifold, e.g., an image manifold. The intermediate values between the initial latent space and the image manifold can be interpreted as noisy images, with the amount of noise determined by the forward diffusion process noise schedule. We utilize this interpretation to present Boomerang, an approach for local sampling of image manifolds. As implied by its name, Boomerang local sampling involves adding noise to an input image, moving it closer to the latent space, and then mapping it back to the image manifold through a partial reverse diffusion process. Thus, Boomerang generates images on the manifold that are ``similar,'' but nonidentical, to the original input image. We can control the proximity of the generated images to the original by adjusting the amount of noise added. Furthermore, due to the stochastic nature of the reverse diffusion process in Boomerang, the generated images display a certain degree of stochasticity, allowing us to obtain local samples from the manifold without encountering any duplicates. Boomerang offers the flexibility to work seamlessly with any pretrained diffusion model, such as Stable Diffusion, without necessitating any adjustments to the reverse diffusion process. We present three applications for Boomerang. First, we provide a framework for constructing privacy-preserving datasets having controllable degrees of anonymity. Second, we show that using Boomerang for data augmentation increases generalization performance and outperforms state-of-the-art synthetic data augmentation. Lastly, we introduce a perceptual image enhancement framework, which enables resolution enhancement.
연구 동기 및 목표
- 모델 미세조정이나 아키텍처 변경 없이 사전 훈련된 확산 모델을 사용하여 이미지 매니폴드에서 국소적 샘플링을 가능하게 하기.
- 원본 이미지의 분포에 가까운, 서로 다른 비동일한 변형을 생성할 수 있는 방법을 제공하기.
- 개인 정보 보호 데이터 생성, 데이터 증강, 고해상도 이미지 초해상도 처리와 같은 실용적 응용을 지원하기.
- 확산 모델의 역동적 특성을 활용하여 초해상도와 같은 역문제의 타당성을 탐색하기.
- 다양한 확산 모델 아키텍처와 노이즈 스케줄링 방식에서의 방법의 견고성과 제어 가능성 평가하기.
제안 방법
- 입력 이미지의 잠재 표현에 정방향 확산 스케줄을 사용하여 노이즈를 추가하여 잠재 공간으로 이동시키기.
- 노이즈 수준에 해당하는 선택된 역방향 단계 $t_{\text{Boomerang}}$ 에서 확산 과정을 뒤집어 이미지 매니폴드로 매핑하기.
- 노이즈 수준(즉, 역방향 단계 $t_{\text{Boomerang}}$)을 조절하여 원본 이미지와의 유사성 제어하기.
- 확산 모델의 확률적 성질을 활용하여 동일한 입력 이미지의 다수의 서로 다른 실재감 있는 변형 생성하기.
- 고해상도 초해상도 결과의 안정성을 향상시키기 위해 보머랑을 반복 적용하여 캐스케이딩 적용하기.
- 재학습이나 아키텍처 수정 없이 사전 훈련된 확산 모델(예: 스테이블 디퓨전, 패치드 디퓨전) 활용하기.
실험 결과
연구 질문
- RQ1사전 훈련된 확산 모델의 역동적 특성만을 사용하여 이미지 매니폴드에서 국소적 샘플링을 수행할 수 있는가?
- RQ2추가된 노이즈 수준이 원본 이미지에 비해 생성된 이미지의 유사성과 다양성에 어떤 영향을 미치는가?
- RQ3보머랑은 제어 가능한 익명성으로 개인 정보 보호 데이터 생성에 효과적으로 활용될 수 있는가?
- RQ4보머랑은 매니폴드 일致성을 유지하면서 데이터 증강에 얼마나 효과적으로 기여할 수 있는가?
- RQ5보머랑은 미세조정이나 특수 훈련 없이도 고해상도 초해상도(예: 8배)를 달성할 수 있는가?
주요 결과
- 보머랑은 원본 입력에 가까운 다양한 실재감 있는 이미지 변형을 성공적으로 생성하며, 노이즈 수준 $t_{\text{Boomerang}}$ 를 통해 유사성 제어가 가능하다.
- 패치드 디퓨전 모델에서 $t_{\text{Boomerang}} \approx 100$ 으로 설정할 경우 날카움과 충실도 사이의 균형을 잘 이루었으며, 경험적 테스트에서 PSNR를 최대화했다.
- 이 방법은 추가적인 미세조정 없이도 단일 사전 훈련된 확산 모델만으로 8배 초해상도 이미지 생성을 가능하게 한다.
- 캐스케이딩된 보머랑은 초해상도에서의 안정성을 향상시키며, 반복적인 해상도 향상과 사용자 제어 기반의 세부 정보 선택을 가능하게 한다.
- 보머랑의 성능은 노이즈 스케줄링에 따라 달라지며, 노이즈가 이미지 공간에 존재하는 패치드 디퓨전에서는 잘 작동하지만, 잠재 공간에 노이즈가 존재하는 스테이블 디퓨전에서는 낮은 $t_{\text{Boomerang}}$ 에서는 덜 효과적이다.
- 이 방법은 효율적이며 단일 저비용 GPU에서 실행되어 실용적 구현에 접근 가능하다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.