Skip to main content
QUICK REVIEW

[论文解读] From Infinite to Finite Programs: Explicit Error Bounds with Applications to Approximate Dynamic Programming

Peyman Mohajerin Esfahani, Tobias Sutter|arXiv (Cornell University)|Jan 23, 2017
Risk and Portfolio Optimization参考文献 35被引用 4
一句话总结

本文提出了一种用于马尔可夫决策过程(MDPs)中出现的无限维线性规划(LP)问题的新型近似框架,通过使用正则化有限凸规划,提供显式的先验和后验误差界。结合随机优化与一阶方法,该方法为平均成本和折扣成本最优控制问题建立了紧密的性能保证,实现了在不可数状态与动作空间中的可处理、可认证的解决方案。

ABSTRACT

We consider linear programming (LP) problems in infinite dimensional spaces that are in general computationally intractable. Under suitable assumptions, we develop an approximation bridge from the infinite-dimensional LP to tractable finite convex programs in which the performance of the approximation is quantified explicitly. To this end, we adopt the recent developments in two areas of randomized optimization and first order methods, leading to a priori as well as a posterior performance guarantees. We illustrate the generality and implications of our theoretical results in the special case of the long-run average cost and discounted cost optimal control problems for Markov decision processes on Borel spaces. The applicability of the theoretical results is demonstrated through a constrained linear quadratic optimal control problem and a fisheries management problem.

研究动机与目标

  • 解决最优控制与动态规划中无限维LP的计算不可行性问题。
  • 为具有不可数状态与动作空间的MDP开发一种具有显式误差界的构造性近似方案。
  • 通过统一的正则化有限规划框架,整合并扩展现有针对折扣成本与平均成本问题的方法。
  • 利用随机优化与凸优化技术,提供先验与后验误差保证。
  • 通过带约束的LQR与渔业管理问题展示适用性,并进行严格的性能量化。

提出的方法

  • 将无限时域MDP问题表述为在测度上的无限维LP,利用对偶性与矩问题。
  • 通过将决策变量限制在有限维子空间并使用随机方法采样约束,引入正则化半无限规划。
  • 对对偶变量施加范数约束(正则化项)以限制优化器范数,并推导显式误差界。
  • 应用一阶方法求解所得有限凸规划,并保证收敛性。
  • 通过下确界条件与算子范数推导先验界,通过对偶间隙估计获得后验界。
  • 利用强对偶性与约束采样,确保鲁棒性与有限样本性能保证。

实验结果

研究问题

  • RQ1能否为MDP中无限维LP的有限维近似推导出显式的先验与后验误差界?
  • RQ2如何构建一个正则化有限规划,以确保在不可数MDP中优化器有界且误差控制紧密?
  • RQ3范数约束(正则化项)在稳定对偶解并实现误差量化方面起什么作用?
  • RQ4与现有渐近保证相比,所提出的误差界在实际可用性与计算可处理性方面表现如何?
  • RQ5该框架能否以统一的理论处理方式应用于折扣成本与平均成本MDP?

主要发现

  • 所提出的正则化有限规划实现了显式的先验误差界,其规模与采样约束数量及问题的算子范数相关。
  • 对于折扣成本问题,对偶优化器范数满足 $ \|y^\star\|_{\mathrm{W}} \leq \frac{\theta_{\mathcal{P}} + (1-\tau)^{-1}\|\psi\|_{\infty}}{(1-\tau)\theta_{\mathcal{P}} - \|\psi\|_{\mathrm{L}}} $,确保了稳定性。
  • 后验误差界通过对偶间隙推导,数值结果表明其收敛于真实最优值的紧密范围内。
  • 该框架成功统一处理了长期平均成本与折扣成本问题,具有统一的理论结构。
  • 在LQR与渔业管理示例中,算法实现了接近最优的性能,并具备可认证的误差界,验证了其实际适用性。
  • 该方法为渐近方案提供了构造性替代,首次在有限样本下提供了性能保证。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。