Skip to main content
QUICK REVIEW

[논문 리뷰] Perception Prioritized Training of Diffusion Models

Jooyoung Choi, Jungbeom Lee|arXiv (Cornell University)|2022. 04. 01.
Generative Adversarial Networks and Image Synthesis인용 수 7
한 줄 요약

이 논문은 시각적 개념을 인지적으로 풍부하게 학습하는 데에 중요한 노이즈 수준을 우선시하는 간단하면서도 효과적인 재가중 기법인 Perception Prioritized (P2) 가중치를 제안한다. 이미지의 내용이 여전히 구분 가능한 중간 노이즈 수준에 더 높은 손실 가중치를 할당함으로써, 다양한 데이터셋, 아키텍처, 샘플링 전략에서 샘플 품질이 크게 향상되며, CelebA-HQ와 Oxford-Flowers에서 최신 기준(FID) 점수를 달성한다.

ABSTRACT

Diffusion models learn to restore noisy data, which is corrupted with different levels of noise, by optimizing the weighted sum of the corresponding loss terms, i.e., denoising score matching loss. In this paper, we show that restoring data corrupted with certain noise levels offers a proper pretext task for the model to learn rich visual concepts. We propose to prioritize such noise levels over other levels during training, by redesigning the weighting scheme of the objective function. We show that our simple redesign of the weighting scheme significantly improves the performance of diffusion models regardless of the datasets, architectures, and sampling strategies.

연구 동기 및 목표

  • 학습 중에 다양한 노이즈 수준에서 디퓨전 모델이 어떻게 시각적 개념을 학습하는지 조사하기 위해.
  • 디퓨전 모델 학습에서 손실 가중치의 체계적인 이해와 최적화 부족 문제를 해결하기 위해.
  • 아키텍처 변경이나 추가 학습 단계 없이 샘플 품질을 향상시키기 위해.
  • 세부적인, 인지하기 어려운 디테일보다는 시각적으로 정보가 풍부한 노이즈 수준을 우선시하는 가중치 기법을 설계하기 위해.

제안 방법

  • 모델이 각 노이즈 수준에서 학습하는 시각적 개념의 인지적 중요도를 기반으로 디노이징 스코어 매칭 손실을 재가중하며, 이미지의 내용이 여전히 인지 가능성이 있는 중간 노이즈 수준에 더 높은 가중치를 할당한다.
  • P2 가중치 기법은 각 노이즈 수준에서 모델이 어떤 시각적 개념을 학습하는지에 대한 경험적 분석에서 유도되며, 고수준의 인지적으로 풍부한 특징 학습을 지원하는 수준에 더 높은 가중치를 부여한다.
  • 가중치 함수는 중간 노이즈 수준에서 최고점을 가지며, 매우 낮거나 매우 높은 노이즈 수준으로 갈수록 감소하는 조각별 선형 또는 부드러운 함수로 설계된다.
  • 이 방법은 표준 디퓨전 학습 및 샘플링 프로토콜과 호환되며, 모델 아키텍처나 추론 단계에 대한 변경이 필요 없다.
  • 일관된 일반화를 검증하기 위해 여러 데이터셋(CelebA-HQ, FFHQ, Oxford-Flowers), 아키텍처, 샘플링 스케줄에서 평가된다.
Figure 1 : Information removal of a diffusion process. (Left) Perceptual distance of corrupted images as a function of signal-to-noise ratio (SNR). Distances are measured between two noisy images either corrupted from the same image (blue) or different images (orange). We averaged distances measured
Figure 1 : Information removal of a diffusion process. (Left) Perceptual distance of corrupted images as a function of signal-to-noise ratio (SNR). Distances are measured between two noisy images either corrupted from the same image (blue) or different images (orange). We averaged distances measured

실험 결과

연구 질문

  • RQ1디퓨전 과정의 어떤 노이즈 수준이 인지적으로 풍부한 시각적 개념 학습에 가장 기여하는가?
  • RQ2노이즈 수준에 따라 손실 가중치의 분포가 생성된 이미지의 품질에 어떤 영향을 미치는가?
  • RQ3단순하면서도 원리에 기반한 재가중 기법이 아키텍처 수정 없이도 샘플 품질 향상에 기여할 수 있는가?
  • RQ4제안된 가중치 기법은 다양한 데이터셋, 모델 아키텍처, 샘플링 전략에 일반화되는가?

주요 결과

  • P2 가중치 기법은 CelebA-HQ에서 최신 기준 FID 점수 8.92를 달성하여 이전 방법들을 능가한다.
  • Oxford-Flowers 데이터셋에서는 새로운 최신 기준 FID 10.12를 기록하며 기준선 대비 뚜렷한 향상을 이룬다.
  • FFHQ에서는 250개의 샘플링 단계를 사용할 때 FID 10.88을 달성하여 기준선을 초월하며, GAN 기반 모델과 유사한 성능을 보인다.
  • 다양한 샘플링 스케줄에서 일관된 향상이 나타나며, P2 방법은 단지 60단계만 사용할 때에도 기준선을 능가한다.
  • 모델 아키텍처나 샘플링 전략에 관계없이 성능 향상이 이루어지며, 광범위한 일반화 능력을 입증한다.
  • 제거 분석 결과, 내용이 여전히 인지 가능한 중간 노이즈 수준을 우선시하는 것이 균일하거나 표준 가중치 손실 기법보다 더 나은 샘플 품질을 이끌어낸다.
Figure 2 : Stochastic reconstruction. (Left) Illustration of reconstruction, where sample are obtained from full sampling chain. (Right) Reconstructions $\hat{x}_{0}$ with input images $x_{0}$ on the rightmost column and SNR of $x_{t}$ on the bottom. Samples in the 1st, 2nd columns share only the co
Figure 2 : Stochastic reconstruction. (Left) Illustration of reconstruction, where sample are obtained from full sampling chain. (Right) Reconstructions $\hat{x}_{0}$ with input images $x_{0}$ on the rightmost column and SNR of $x_{t}$ on the bottom. Samples in the 1st, 2nd columns share only the co

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.