[Paper Review] A Variational Perspective on Diffusion-Based Generative Models and Score Matching
The paper builds a continuous-time variational framework for diffusion-based generative models and score matching, linking score matching to a lower bound on the likelihood of plug-in reverse SDEs via Feynman-Kac and Girsanov theory.
Discrete-time diffusion-based generative models and score matching methods have shown promising results in modeling high-dimensional image data. Recently, Song et al. (2021) show that diffusion processes that transform data into noise can be reversed via learning the score function, i.e. the gradient of the log-density of the perturbed data. They propose to plug the learned score function into an inverse formula to define a generative diffusion process. Despite the empirical success, a theoretical underpinning of this procedure is still lacking. In this work, we approach the (continuous-time) generative diffusion directly and derive a variational framework for likelihood estimation, which includes continuous-time normalizing flows as a special case, and can be seen as an infinitely deep variational autoencoder. Under this framework, we show that minimizing the score-matching loss is equivalent to maximizing a lower bound of the likelihood of the plug-in reverse SDE proposed by Song et al. (2021), bridging the theoretical gap.
Motivation & Objective
- Motivate and formalize likelihood estimation for continuous-time diffusion processes used in diffusion models.
- Bridge score matching losses with maximum likelihood through a variational ELBO framework.
- Show that minimizing score-matching loss maximizes a lower bound on the marginal likelihood of the plug-in reverse SDE.
- Generalize results to a family of marginal-equivalent plug-in reverse SDEs, including an equivalent ODE as a limiting case.
Proposed method
- Derive a variational framework using Feynman-Kac to represent marginal densities of generative diffusion via expectations.
- Apply Girsanov's change of measure to infer latent Brownian paths and obtain a continuous-time ELBO (CT-ELBO).
- Reparameterize generative and inference SDEs to reveal connections to score functions.
- Prove that the continuous-time ELBO tightens when the inference SDE matches the score of the marginal density.
- Show that the CT-ELBO extends discrete-time ELBO to infinitely deep hierarchies and relates to an infinite-depth VAE viewpoint.
- Discuss computational trade-offs, bias-variance considerations, and debiasing strategies for practical estimation.
Experimental results
Research questions
- RQ1How does minimizing the score-matching loss influence the sampling behavior of the plug-in reverse SDE?
- RQ2Can a variational ELBO framework consistently estimate the marginal likelihood for continuous-time diffusion models?
- RQ3What is the relationship between score matching and maximum likelihood in diffusion-based generative models?
- RQ4Do plug-in reverse SDEs form a continuum whose ELBO is maximized by score matching?
Key findings
- A variational framework yields a continuous-time ELBO that lower-bounds the log marginal density of the plug-in reverse SDE.
- Minimizing the score-matching loss corresponds to maximizing a lower bound on the likelihood of the plug-in reverse SDE, bridging score matching and likelihood estimation.
- The framework extends discrete-time diffusion models to infinite depth, aligning with an infinitely deep hierarchical VAE perspective.
- There exists a family of marginal-equivalent plug-in reverse SDEs, including an equivalent ODE as a limiting case, all sharing the same marginal distributions under certain conditions.
- A detailed comparison reveals trade-offs in computational efficiency, and debiasing strategies improve likelihood estimation in practice on datasets like MNIST and CIFAR-10.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.