Skip to main content
QUICK REVIEW

[논문 리뷰] ScaleLong: Towards More Stable Training of Diffusion Model via Scaling Network Long Skip Connection

Zhongzhan Huang, Pan Zhou|arXiv (Cornell University)|2023. 10. 20.
Advanced Neuroimaging Techniques and Applications인용 수 5
한 줄 요약

이 논문은 U-Net의 장거리 스킵 연결(LSC) 계수를 스케일링하여 확산 모델의 훈련을 안정화하는 ScaleLong 프레임워크를 제안한다. 이론적으로는 큰 LSC 계수 값이 특징과 기울기 진동을 유도함을 보이며, 실증적으로는 다양한 데이터셋과 아키텍처에서 계수 스케일링이 진동을 감소시키고 훈련 속도를 최대 1.5배로 향상시킴을 입증한다.

ABSTRACT

In diffusion models, UNet is the most popular network backbone, since its long skip connects (LSCs) to connect distant network blocks can aggregate long-distant information and alleviate vanishing gradient. Unfortunately, UNet often suffers from unstable training in diffusion models which can be alleviated by scaling its LSC coefficients smaller. However, theoretical understandings of the instability of UNet in diffusion models and also the performance improvement of LSC scaling remain absent yet. To solve this issue, we theoretically show that the coefficients of LSCs in UNet have big effects on the stableness of the forward and backward propagation and robustness of UNet. Specifically, the hidden feature and gradient of UNet at any layer can oscillate and their oscillation ranges are actually large which explains the instability of UNet training. Moreover, UNet is also provably sensitive to perturbed input, and predicts an output distant from the desired output, yielding oscillatory loss and thus oscillatory gradient. Besides, we also observe the theoretical benefits of the LSC coefficient scaling of UNet in the stableness of hidden features and gradient and also robustness. Finally, inspired by our theory, we propose an effective coefficient scaling framework ScaleLong that scales the coefficients of LSC in UNet and better improves the training stability of UNet. Experimental results on four famous datasets show that our methods are superior to stabilize training and yield about 1.5x training acceleration on different diffusion models with UNet or UViT backbones. Code: https://github.com/sail-sg/ScaleLong

연구 동기 및 목표

  • 확산 모델에서 장거리 스킵 연결이 존재함에도 불구하고 U-Net 훈련이 불안정한 이유를 이해하는 것.
  • 특히 $1/\sqrt{2}$-스케일링 기법에 기반한 LSC 계수 스케일링의 안정화 효과에 대한 이론적 설명을 제공하는 것.
  • 확산 모델에서 훈련 안정성과 수렴 속도를 향상시키는 원칙적인 프레임워크를 개발하는 것.
  • 실증적으로 LSC 계수 스케일링이 특징과 기울기 진동을 감소시키고 훈련 속도를 향상시킴을 검증하는 것.

제안 방법

  • U-Net의 은닉 특징과 기울기 노름에 대한 이론적 경계를 유도하여, 진동 범위가 $\mathcal{O}(m\sum_{i=1}^{N}\kappa_i^2)$로 스케일링되며, 여기서 $\kappa_i$는 LSC 계수임을 보임.
  • U-Net의 입력 편향에 대한 내성적 안정성은 $\mathcal{O}(\sum_{i=1}^{N}\kappa_i M_0^i)$로 경계지며, $M_0 > 1$일 경우 불안정성 위험이 있음을 증명함.
  • 진동 범위를 줄이고 훈련 안정성을 향상시키기 위해 LSC에 적응형 계수 스케일링을 적용하는 ScaleLong 프레임워크를 제안함.
  • CIFAR10, CelebA, ImageNet64에서 특징 시각화와 기울기 추적을 통해 이론적 경계를 실증적으로 검증함.
  • ScaleLong 내에서 $1/\sqrt{2}$-스케일링 기반보다 더 효과적인 두 가지 신규이지만 효과적인 계수 스케일링 방법을 도입함.
  • 제거 실험을 통해 $\kappa > 1$은 특히 작은 배치 크기에서 불안정성을 악화시키며, $\kappa \in (0,1)$일 경우 안정성을 보장함을 입증함.

실험 결과

연구 질문

  • RQ1확산 모델에서 장거리 스킵 연결이 존재함에도 불구하고 U-Net이 불안정한 훈련을 보이는 이유는 무엇인가?
  • RQ2장거리 스킵 연결의 계수는 U-Net의 전방 및 역전파 안정성에 어떻게 영향을 미치는가?
  • RQ3$1/\sqrt{2}$-스케일링이 LSC 계수에 미치는 안정화 효과의 이론적 근거는 무엇인가?
  • RQ4경험적 히우리즘을 초월해 원칙적인 계수 스케일링 프레임워크가 훈련 안정성과 수렴 속도를 향상시킬 수 있는가?
  • RQ5U-Net의 입력 편향에 대한 내성적 안정성은 LSC 계수에 어떻게 의존하는가?

주요 결과

  • LSC 계수 $\kappa_i = 1$일 경우 U-Net의 은닉 특징과 기울기의 진동 범위는 $\mathcal{O}(mN)$이며, 이는 훈련 중 관찰된 불안정성의 원인을 설명한다.
  • LSC 계수를 $1/\sqrt{2}$로 스케일링하면 진동 범위가 감소하고 훈련 안정성이 향상되지만, 불안정성이 완전히 제거되지는 않는다.
  • 이론적 경계에 따르면 기울기 크기는 $\mathcal{O}(m\sum_{i=1}^{N}\kappa_i^2)$로 상한이 있으며, $\kappa_i = 1$일 경우 이 값이 크므로 매개변수 갱신이 불안정해짐.
  • U-Net의 입력 편향에 대한 내성적 안정성은 $\mathcal{O}(\sum_{i=1}^{N}\kappa_i M_0^i)$로 경계지며, $M_0 > 1$일 경우 작은 입력 변화에 민감함을 나타냄.
  • ScaleLong은 특징과 기울기 진동을 감소시켜 UNet 또는 UViT 백본을 가진 여러 확산 모델에서 훈련 속도를 최대 1.5배로 향상시킴.
  • 실험 결과 $\kappa > 1$은 특히 작은 배치 크기에서 빠른 훈련 붕괴를 유도하며, 이는 이론적 도메인 제약 조건인 $\kappa \in (0,1)$을 검증함.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.