[Paper Review] Benchmarking Uncertainty Quantification of Plug-and-Play Diffusion Priors for Inverse Problems Solving
The paper benchmarks uncertainty quantification (UQ) of plug-and-play diffusion prior (PnPDP) solvers for inverse problems, showing that similar reconstruction quality can hide vastly different posterior uncertainty, and proposes a UQ-driven taxonomy and diagnostic framework.
Plug-and-play diffusion priors (PnPDP) have become a powerful paradigm for solving inverse problems in scientific and engineering domains. Yet, current evaluations of reconstruction quality emphasize point-estimate accuracy metrics on a single sample, which do not reflect the stochastic nature of PnPDP solvers and the intrinsic uncertainty of inverse problems, critical for scientific tasks. This creates a fundamental mismatch: in inverse problems, the desired output is typically a posterior distribution and most PnPDP solvers induce a distribution over reconstructions, but existing benchmarks only evaluate a single reconstruction, ignoring distributional characterization such as uncertainty. To address this gap, we conduct a systematic study to benchmark the uncertainty quantification (UQ) of existing diffusion inverse solvers. Specifically, we design a rigorous toy model simulation to evaluate the uncertainty behavior of various PnPDP solvers, and propose a UQ-driven categorization. Through extensive experiments on toy simulations and diverse real-world scientific inverse problems, we observe uncertainty behaviors consistent with our taxonomy and theoretical justification, providing new insights for evaluating and understanding the uncertainty for PnPDPs.
Motivation & Objective
- Motivate the need for uncertainty-aware evaluation of PnPDP solvers in ill-posed inverse problems.
- Propose a UQ-driven categorization of PnPDP methods based on their ability to approximate the Bayesian posterior.
- Develop toy-model diagnostics to quantify calibration of uncertainty and compare methods.
- Demonstrate consistent uncertainty behaviors across real-world scientific inverse problems.
- Provide recommendations for evaluating and understanding uncertainty in diffusion-based inverse solvers.
Proposed method
- Define the posterior target p(x|y) and distinguish posterior-targeting, heuristic, and MAP-like PnPDP solvers.
- Introduce empirical posterior variance from repeated samples as a proxy for solver-induced uncertainty.
- Design toy experiments with known ground-truth posterior to calibrate AU and EU and validate UQ metrics.
- Evaluate real-data inverse problems (linear scattering, sparse-sampling MRI, sparse-view CT) across multiple PnPDP methods.
- Provide a UQ-driven taxonomy linking methods to their posterior-targeting capabilities and theoretical guarantees.

Experimental results
Research questions
- RQ1Can stochastic PnPDP solvers recover the posterior p(x|y) and its uncertainty under ill-posed forward models?
- RQ2How do different PnPDP solvers compare in terms of calibrated uncertainty, not just reconstruction accuracy?
- RQ3Does a UQ-driven categorization reflect observed uncertainty behaviors across toy and real-data tasks?
- RQ4How does measurement sparsity or out-of-distribution data affect solver uncertainty?
- RQ5What are the limitations and practical gaps between asymptotic guarantees and real-world implementations of posterior-targeting solvers?
Key findings
- Uncertainty behavior of PnPDP solvers can vary widely even when reconstruction accuracy is similar.
- Posterior-targeting solvers (e.g., MCG-Diff, FPS-SMC, PnPDM) show calibrated uncertainty in toy and some real-data tests, but can still exhibit bias or degeneracy in practice.
- MAP-like solvers (e.g., REDDiff) yield near-zero variance, consistent with point estimation.
- Heuristic solvers display diverse uncertainty patterns and can fail to calibrate uncertainty under certain forward models.
- Uncertainty generally increases with measurement sparsity, and accuracy and uncertainty are complementary assessments.
- A UQ-driven taxonomy aligns with observed behaviors and complements algorithm-structure classifications.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.