Skip to main content
QUICK REVIEW

[論文レビュー] Diffusion Probabilistic Model Made Slim

Xingyi Yang, Daquan Zhou|arXiv (Cornell University)|Nov 27, 2022
Generative Adversarial Networks and Image Synthesis被引用数 4
ひとこと要約

本論文は、ウェーブレットゲーティングとスペクトル認識 distillation を通じて、コンactアーキテクチャにおける高周波成分の復元を向上させる軽量な拡散モデルであるSpectral Diffusion (SD) を提案する。逆ノイズ除去中に周波数ダイナミクスをモデル化し、逆数重み付き損失を用いて高周波成分の知識蒸留を行うことで、標準的な潜在拡散モデルと比較して 8–18× 小さなモデルサイズと 2–5× の高速な推論を達成しつつ、忠実度の低下を最小限に抑える。

ABSTRACT

Despite the recent visually-pleasing results achieved, the massive computational cost has been a long-standing flaw for diffusion probabilistic models (DPMs), which, in turn, greatly limits their applications on resource-limited platforms. Prior methods towards efficient DPM, however, have largely focused on accelerating the testing yet overlooked their huge complexity and sizes. In this paper, we make a dedicated attempt to lighten DPM while striving to preserve its favourable performance. We start by training a small-sized latent diffusion model (LDM) from scratch, but observe a significant fidelity drop in the synthetic images. Through a thorough assessment, we find that DPM is intrinsically biased against high-frequency generation, and learns to recover different frequency components at different time-steps. These properties make compact networks unable to represent frequency dynamics with accurate high-frequency estimation. Towards this end, we introduce a customized design for slim DPM, which we term as Spectral Diffusion (SD), for light-weight image synthesis. SD incorporates wavelet gating in its architecture to enable frequency dynamic feature extraction at every reverse steps, and conducts spectrum-aware distillation to promote high-frequency recovery by inverse weighting the objective based on spectrum magni tudes. Experimental results demonstrate that, SD achieves 8-18x computational complexity reduction as compared to the latent diffusion models on a series of conditional and unconditional image generation tasks while retaining competitive image fidelity.

研究の動機と目的

  • コンact拡散モデルが高周波成分の画像詳細を保持できないという課題に対処すること。
  • アーキテクチャの簡略化にもかかわらず、テクスチャやエッジの回復に劣る小さな拡散モデルの理由を調査すること。
  • 逆ノイズ除去プロセスにおいて周波数ダイナミクスとバイアスを明示的にモデル化する手法を開発し、コンactモデルの性能を向上させること。
  • 条件付きおよび無条件生成タスクにおいて、画像品質を損なわずに顕著なモデル圧縮と推論速度向上を達成すること。

提案手法

  • U-Netのスキップ接続にウェーブレットゲーティングを導入し、各逆ノイズ除去ステップで高周波、低周波、バンドパス成分を動的に調整する。
  • ウェーブレット変換された特徴量に学習可能なゲーティング関数を適用することで、パrameter数を増加させずに適応的周波数応答を実現する。
  • 周波数の大きさに反比例する重みを付けて損失を重みづけることで、スペクトル認識 distillation を実装し、高周波成分の回復を強調する。
  • 事前学習済みの教師モデルを用いて小さな学生モデルを訓練し、損失成分を空間領域と周波数領域に分離して最適化をターゲット化する。
  • 離散フーリエ変換(DFT)分析を用いて、生成画像における周波数バイアスとその変化を診断し、設計選択の妥当性を検証する。
  • ウェーブレットゲーティングとスペクトル認識 distillation を統合した訓練目的関数を構築し、スリムモデルにおける高周波成分の保持を実現する。
Figure 1 : (1) Visualizing the frequency domain gap among generated images with the full DPM [ 43 ] , Lite DPM and our SD on FFHQ [ 26 ] dataset. Lite-DPM is unable to recover fine-grained textures, while SD can produce sharp edges and realistic patterns. (2) Model size, Multiply-Add cumulation (MAC
Figure 1 : (1) Visualizing the frequency domain gap among generated images with the full DPM [ 43 ] , Lite DPM and our SD on FFHQ [ 26 ] dataset. Lite-DPM is unable to recover fine-grained textures, while SD can produce sharp edges and realistic patterns. (2) Model size, Multiply-Add cumulation (MAC

実験結果

リサーチクエスチョン

  • RQ1なぜ小さな拡散モデルはテクスチャーやエッジのような高周波成分を回復できないのか?
  • RQ2標準的な拡散モデルにおいて、逆ノイズ除去ステップの進行に伴い、出力の周波数コンテンツはどのように変化するのか?
  • RQ3ノイズ除去プロセスにおける周波数バイアスが、小さなモデルの微細なディテール再構成能力にどの程度悪影響を及えるのか?
  • RQ4ウェーブレットベースのゲーティングは、コンact拡散モデルにおける周波数に配慮した特徴表現を改善できるか?
  • RQ5知識蒸留における周波数成分の逆数重み付けは、小さなモデルにおける高周波成分の回復を向上させることができるか?

主な発見

  • Spectral Diffusion は、標準的な潜在拡散モデルと比較して、8–18× のモデルサイズ削減と 2–5× の推論時間短縮を達成しながら、競争力のある画像品質を維持する。
  • アブレーションスタディでは、ウェーブレットゲーティングを除去すると FFHQ で FID が 10.5 から 12.4 に上昇し、画像ディテールの保持に果たす重要性が裏付けられる。
  • スペクトル認識 distillation は 1.8 FID ポints の改善をもたらし、空間のみの distillation より顕著に優れている。
  • 77.6M パラメータで、MS-COCO のテキスト到画像生成タスクで FID 18.43 を達成し、ベースライン LDM と比較して 18.7× 小さなモデルサイズである。
  • 可視化結果から、SD は髪の毛や建築パターンのような微細なテクスチャを回復している一方、ベースラインモデルは滑らかで詳細のない出力を生成することが確認された。
  • ゲーティング関数の値はノイズ除去ステップに伴い変化する:後期のアップサンプリング段階で高周波成分が強化され、理論的な周波数進化と整合的である。
Figure 2 : Illustration of the Frequency Evolution and Bias for Diffusion Models. In the reverse process, the optimal filters recover low-frequency components first and add on the details at the end. The predicted score functions may be incorrect for rare patterns, thus failing to recover complex an
Figure 2 : Illustration of the Frequency Evolution and Bias for Diffusion Models. In the reverse process, the optimal filters recover low-frequency components first and add on the details at the end. The predicted score functions may be incorrect for rare patterns, thus failing to recover complex an

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。