Skip to main content
QUICK REVIEW

[Paper Review] From Infinite to Finite Programs: Explicit Error Bounds with Applications to Approximate Dynamic Programming

Peyman Mohajerin Esfahani, Tobias Sutter|arXiv (Cornell University)|Jan 23, 2017
Risk and Portfolio Optimization35 references4 citations
TL;DR

This paper proposes a novel approximation framework for infinite-dimensional linear programs (LPs) arising in Markov decision processes (MDPs), using a regularized finite convex program to provide explicit a priori and a posteriori error bounds. By combining randomized optimization and first-order methods, it establishes tight performance guarantees for both discounted and average cost optimal control problems, enabling tractable, certified solutions in uncountable state and action spaces.

ABSTRACT

We consider linear programming (LP) problems in infinite dimensional spaces that are in general computationally intractable. Under suitable assumptions, we develop an approximation bridge from the infinite-dimensional LP to tractable finite convex programs in which the performance of the approximation is quantified explicitly. To this end, we adopt the recent developments in two areas of randomized optimization and first order methods, leading to a priori as well as a posterior performance guarantees. We illustrate the generality and implications of our theoretical results in the special case of the long-run average cost and discounted cost optimal control problems for Markov decision processes on Borel spaces. The applicability of the theoretical results is demonstrated through a constrained linear quadratic optimal control problem and a fisheries management problem.

Motivation & Objective

  • To address the computational intractability of infinite-dimensional LPs in optimal control and dynamic programming.
  • To develop a constructive approximation scheme with explicit error bounds for MDPs with uncountable state and action spaces.
  • To unify and extend existing approaches for discounted and average cost problems using a common regularized finite program framework.
  • To provide both a priori and a posteriori error guarantees using randomized and convex optimization techniques.
  • To demonstrate applicability through constrained LQR and fisheries management problems with rigorous performance quantification.

Proposed method

  • Formulates the infinite-horizon MDP problem as an infinite-dimensional LP over measures, leveraging duality and moment problems.
  • Introduces a regularized semi-infinite program by restricting decision variables to a finite-dimensional subspace and sampling constraints via randomized methods.
  • Employs a norm constraint (regularizer) on dual variables to bound optimizer norms and derive explicit error bounds.
  • Applies first-order methods to solve the resulting finite convex programs with convergence guarantees.
  • Derives a priori bounds using inf-sup conditions and operator norms, and a posteriori bounds via duality gap estimation.
  • Uses strong duality and constraint sampling to ensure robustness and finite-sample performance guarantees.

Experimental results

Research questions

  • RQ1Can explicit a priori and a posteriori error bounds be derived for finite-dimensional approximations of infinite-dimensional LPs in MDPs?
  • RQ2How can a regularized finite program be constructed to ensure bounded optimizers and tight error control in uncountable MDPs?
  • RQ3What is the role of the norm constraint (regularizer) in stabilizing the dual solution and enabling error quantification?
  • RQ4How do the proposed bounds compare to existing asymptotic guarantees in terms of practical usability and computational tractability?
  • RQ5Can the framework be applied to both discounted and average cost MDPs with a unified theoretical treatment?

Key findings

  • The proposed regularized finite program achieves explicit a priori error bounds that scale with the number of sampled constraints and the problem's operator norm.
  • For the discounted cost problem, the dual optimizer norm is bounded by $ \|y^\star\|_{\mathrm{W}} \leq \frac{\theta_{\mathcal{P}} + (1-\tau)^{-1}\|\psi\|_{\infty}}{(1-\tau)\theta_{\mathcal{P}} - \|\psi\|_{\mathrm{L}}} $, ensuring stability.
  • A posteriori error bounds are derived from the duality gap, with numerical results showing convergence to the true optimal value within tight bounds.
  • The framework successfully handles both long-run average cost and discounted cost problems under a unified theoretical structure.
  • In the LQR and fisheries management examples, the algorithm achieves near-optimal performance with certified error bounds, validating practical applicability.
  • The method provides a constructive alternative to asymptotic schemes, offering finite-sample performance guarantees where none existed before.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.