[Paper Review] Maximum likelihood estimation of regularisation parameters in high-dimensional inverse problems: an empirical Bayesian approach. Part II: Theoretical Analysis
This paper provides a rigorous theoretical analysis of a stochastic approximation proximal gradient method for maximum marginal likelihood estimation of regularization parameters in high-dimensional inverse problems. It establishes almost sure convergence of the algorithm under mild, verifiable conditions, with explicit non-asymptotic convergence rates, using inexact proximal Markov chain Monte Carlo samplers—specifically proximal Langevin algorithms—enabling scalable and theoretically grounded optimization in high dimensions.
This paper presents a detailed theoretical analysis of the three stochastic approximation proximal gradient algorithms proposed in our companion paper [49] to set regularization parameters by marginal maximum likelihood estimation. We prove the convergence of a more general stochastic approximation scheme that includes the three algorithms of [49] as special cases. This includes asymptotic and non-asymptotic convergence results with natural and easily verifiable conditions, as well as explicit bounds on the convergence rates. Importantly, the theory is also general in that it can be applied to other intractable optimisation problems. A main novelty of the work is that the stochastic gradient estimates of our scheme are constructed from inexact proximal Markov chain Monte Carlo samplers. This allows the use of samplers that scale efficiently to large problems and for which we have precise theoretical guarantees.
Motivation & Objective
- To establish almost sure convergence of a general stochastic approximation scheme for marginal likelihood optimization in high-dimensional inverse problems.
- To provide non-asymptotic convergence rates with easily verifiable conditions for the proposed algorithm.
- To analyze the convergence of a method that uses inexact proximal MCMC samplers (e.g., MYULA) for gradient estimation in high-dimensional settings.
- To generalize the theoretical framework to apply beyond the three specific algorithms in the companion paper, covering them as special cases.
- To ensure theoretical guarantees for optimization in ill-posed or ill-conditioned imaging problems where regularization parameter selection is critical.
Proposed method
- Proposes a general stochastic approximation scheme of the form: θn+1 = ΠΘ[θn − δn+1/mn ∑(g(Xn,k) − g(¯Xn,k))], where (Xn,k) and (¯Xn,k) are inexact MCMC samplers targeting the posterior and prior distributions.
- Employs proximal Langevin samplers (e.g., MYULA) to handle non-differentiable regularizers like total variation or ℓ1-norms, ensuring scalability in high dimensions.
- Uses inexact MCMC samplers with controlled bias and variance, where the samplers are based on generalized Unadjusted Langevin Algorithms (ULA) with proximal steps.
- Applies Fisher’s identity to express the gradient of the log-marginal likelihood as a difference of expectations, which are estimated via MCMC samples.
- Establishes convergence using a framework from stochastic approximation theory, with conditions on stepsize sequences and batch sizes.
- Derives explicit bounds on convergence rates by analyzing the Kullback-Leibler divergence between transition kernels and using generalized Pinsker inequalities and Lyapunov functions.
Experimental results
Research questions
- RQ1Under what conditions does the stochastic approximation scheme converge almost surely to a solution of the marginal likelihood maximization problem?
- RQ2What are the non-asymptotic convergence rates of the algorithm, and how do they depend on the stepsize and batch size sequences?
- RQ3How does the use of inexact proximal MCMC samplers affect the convergence properties of the optimization scheme?
- RQ4Can the theoretical framework be extended to cover multiple algorithms from the companion paper as special cases?
- RQ5What are the minimal and verifiable assumptions required for convergence when using proximal Langevin samplers in high-dimensional, non-differentiable settings?
Key findings
- The proposed stochastic approximation scheme converges almost surely to a solution of the marginal likelihood maximization problem under mild integrability and regularity conditions.
- Non-asymptotic convergence rates are established with explicit bounds depending on the stepsize sequence (δn), batch size (mn), and the mixing properties of the MCMC samplers.
- The convergence rate is shown to be O(1/√n) under appropriate choices of stepsize and batch size, with explicit dependence on the spectral gap and log-Sobolev constants of the target distributions.
- The use of inexact proximal MCMC samplers (e.g., MYULA) is theoretically justified, with bounds on the bias and variance of the gradient estimates that decay with the number of MCMC iterations.
- The theory applies to a broad class of imaging problems with non-differentiable regularizers (e.g., TV, ℓ1), enabling scalable and robust parameter estimation in high-dimensional settings.
- Theoretical guarantees are established for both the standard and proximal versions of the Unadjusted Langevin Algorithm, with explicit dependence on the stepsize and regularization parameters.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.