Skip to main content
QUICK REVIEW

[論文レビュー] Boomerang: Local sampling on image manifolds using diffusion models

Lorenzo Luzi, P. Mayer|arXiv (Cornell University)|Oct 21, 2022
Generative Adversarial Networks and Image Synthesis被引用数 10
ひとこと要約

Boomerang は、入力画像に制御されたノイズを追加した後、前向き拡散プロセスを逆方向に実行することで、事前学習済みの拡散モデルを用いて画像多様体上のローカルサンプリングを可能にする。これにより、元の画像に近い多様で現実的なバリエーションが生成される。モデルの微調整なしに、類似性と確率的多様性を制御可能であり、プライバシー保護型データ生成、データ拡張、1枚のGPU上での8倍スーパーレゾリューションへの応用が可能である。

ABSTRACT

The inference stage of diffusion models can be seen as running a reverse-time diffusion stochastic differential equation, where samples from a Gaussian latent distribution are transformed into samples from a target distribution that usually reside on a low-dimensional manifold, e.g., an image manifold. The intermediate values between the initial latent space and the image manifold can be interpreted as noisy images, with the amount of noise determined by the forward diffusion process noise schedule. We utilize this interpretation to present Boomerang, an approach for local sampling of image manifolds. As implied by its name, Boomerang local sampling involves adding noise to an input image, moving it closer to the latent space, and then mapping it back to the image manifold through a partial reverse diffusion process. Thus, Boomerang generates images on the manifold that are ``similar,'' but nonidentical, to the original input image. We can control the proximity of the generated images to the original by adjusting the amount of noise added. Furthermore, due to the stochastic nature of the reverse diffusion process in Boomerang, the generated images display a certain degree of stochasticity, allowing us to obtain local samples from the manifold without encountering any duplicates. Boomerang offers the flexibility to work seamlessly with any pretrained diffusion model, such as Stable Diffusion, without necessitating any adjustments to the reverse diffusion process. We present three applications for Boomerang. First, we provide a framework for constructing privacy-preserving datasets having controllable degrees of anonymity. Second, we show that using Boomerang for data augmentation increases generalization performance and outperforms state-of-the-art synthetic data augmentation. Lastly, we introduce a perceptual image enhancement framework, which enables resolution enhancement.

研究の動機と目的

  • 事前学習済みの拡散モデルを用いて、微調整やアーキテクチャの変更なしに、画像多様体上のローカルサンプリングを可能にすること。
  • 元の画像の分布に近く、非同一の多様なバリエーションを生成する手法を提供すること。
  • プライバシー保護型データ合成、データ拡張、高スケールの画像スーパーレゾリューションなどの実用的応用を支援すること。
  • 拡散モデルの逆方向ダイナミクスを用いた逆問題(例:スーパーレゾリューション)の可能性を検討すること。
  • 異なる拡散モデルアーキテクチャおよびノイズスケジューリング方式において、本手法のロバストネスと制御可能性を評価すること。

提案手法

  • 入力画像の潜在表現に、前向き拡散スケジュールに従ってノイズを追加し、潜在空間へと移動させる。
  • ノイズレベルに対応する所定の逆方向ステップ $t_{\text{Boomerang}}$ から拡散プロセスを逆方向に実行し、画像多様体へとマッピングする。
  • ノイズレベル(すなわち、逆方向ステップ $t_{\text{Boomerang}}$)を調整することで、元の画像との類似性を制御する。
  • 拡散モデルの確率的性質を活用し、同一の入力画像から複数の異なる現実的なバリエーションを生成する。
  • 繰り返しBoomerangを適用することでカスケード化し、高スケールのスーパーレゾリューション結果の安定性を向上させ、段階的に解像度を向上させる。
  • 再トレーニングやアーキテクチャの変更なしに、事前学習済みの拡散モデル(例:Stable Diffusion, Patched Diffusion)を活用する。

実験結果

リサーチクエスチョン

  • RQ1事前学習済みの拡散モデルの逆方向ダイナミクスのみを用いて、画像多様体上のローカルサンプリングが可能か?
  • RQ2追加されるノイズのレベルが、生成画像の類似性と多様性に与える影響は何か?
  • RQ3Boomerang は、制御可能な匿名性を備えたプライバシー保護型データ生成に効果的に利用可能か?
  • RQ4Boomerang は、多様体の一貫性を保ちながら、どの程度データ拡張を強化できるか?
  • RQ5Boomerang は、微調整や特別なトレーニングなしに、高スケールの画像スーパーレゾリューション(例:8倍)を達成できるか?

主な発見

  • Boomerang は、元の入力画像に近く、多様で現実的な画像バリエーションを効果的に生成でき、類似性はノイズレベル $t_{\text{Boomerang}}$ で制御可能である。
  • Patched Diffusion モデルにおいて $t_{\text{Boomerang}} \approx 100$ に設定した場合、実験的評価でシャープネスと忠実度のバランスが良く、PSNRが最大値に達した。
  • 本手法により、微調整や再トレーニングなしに、1つの事前学習済み拡散モデルのみを用いて8倍の画像スーパーレゾリューションが実現可能である。
  • カスケード化されたBoomerangは、スーパーレゾリューションにおける安定性を向上させ、反復的な解像度向上とユーザーが制御可能な詳細選択を可能にする。
  • 本手法の性能はノイズスケジューリングに依存する:Patched Diffusion(画像空間でのノイズ)では良好に動作するが、Stable Diffusion(潜在空間でのノイズ)では $t_{\text{Boomerang}}$ が低い場合にやや効果が低い。
  • 本アプローチは効率的であり、1枚の安価なGPU上で動作するため、実用的展開に適している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。