Skip to main content
QUICK REVIEW

[Paper Review] A Survey on Generative Diffusion Model

Hanqun Cao, Cheng Tan|arXiv (Cornell University)|Sep 6, 2022
Creativity in Education and NeurosciencePsychology66 citations
TL;DR

This survey comprehensively analyzes diffusion models across fundamentals, algorithmic improvements, and applications, summarizing forward/reverse formulations, DDPM/score-SDE frameworks, and conditional generation, plus advances in sampling, forward-process design, likelihood optimization, and distribution bridging.

ABSTRACT

Deep generative models have unlocked another profound realm of human creativity. By capturing and generalizing patterns within data, we have entered the epoch of all-encompassing Artificial Intelligence for General Creativity (AIGC). Notably, diffusion models, recognized as one of the paramount generative models, materialize human ideation into tangible instances across diverse domains, encompassing imagery, text, speech, biology, and healthcare. To provide advanced and comprehensive insights into diffusion, this survey comprehensively elucidates its developmental trajectory and future directions from three distinct angles: the fundamental formulation of diffusion, algorithmic enhancements, and the manifold applications of diffusion. Each layer is meticulously explored to offer a profound comprehension of its evolution. Structured and summarized approaches are presented in https://github.com/chq1155/A-Survey-on-Generative-Diffusion-Model.

Motivation & Objective

  • Explain the fundamental diffusion model formulation and its theoretical underpinnings.
  • Summarize key algorithmic improvements and practical techniques boosting diffusion model performance.
  • Survey diverse applications of diffusion models across domains including vision, language, healthcare, and science.
  • Provide taxonomy of advancements and future directions for diffusion models.

Proposed method

  • Explain forward and reverse diffusion processes and transition kernels.
  • Describe DDPM and score-based SDE formalisms and their training objectives.
  • Discuss conditional diffusion probabilistic models and guidance techniques.
  • Present taxonomy of improvements: sampling acceleration, diffusion design, likelihood optimization, and bridging distributions.
  • Outline latent-space diffusion and non-Euclidean/different data space extensions.

Experimental results

Research questions

  • RQ1What are the core theoretical formulations of diffusion models and their connections to DDPM and SDE frameworks?
  • RQ2What algorithmic strategies have been proposed to accelerate sampling and improve training and likelihood objectives?
  • RQ3How are diffusion models extended to conditional generation and non-traditional data domains (latent spaces, discrete spaces, manifolds, graphs)?
  • RQ4What are the main applications and domains where diffusion models have been successfully applied, and what future directions are identified?

Key findings

  • Diffusion models offer stable training and high-quality generation compared to VAEs, EBMs, and NFs, but require slower sampling due to iterative denoising (forward to prior to data).
  • DDPM and score-based SDE formulations provide complementary discrete and continuous frameworks with denoising score-matching objectives.
  • Multiple algorithmic improvements are categorized: sampling acceleration, diffusion process design, likelihood optimization, and bridging distributions.
  • Conditional diffusion models enable generation conditioned on labels or text using guidance techniques like classifier-free guidance.
  • Latent-space and non-Euclidean extensions broaden diffusion applicability to images, language, graphs, molecules, and manifolds.
  • The survey outlines future directions and connections to other diffusion surveys, highlighting limitations and potential research avenues.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.