Skip to main content
QUICK REVIEW

[论文解读] Stabilizing Value Iteration with and without Approximation Errors

Ali Heydari|arXiv (Cornell University)|Dec 17, 2014
Adaptive Dynamic Programming Control参考文献 27被引用 7
一句话总结

该论文为具有连续状态和动作空间的离散时间最优控制中的值迭代(VI)在存在和不存在近似误差的情况下建立了理论稳定条件。它证明了收敛到最优解,并提供了显式、可验证的吸引域估计,确保在线学习过程中即使存在函数近似误差也能保持系统稳定。

ABSTRACT

Adaptive optimal control using value iteration (VI) initiated from a stabilizing policy is theoretically analyzed in various aspects including the continuity of the result, the stability of the system operated using any single/constant resulting control policy, the stability of the system operated using the evolving/time-varying control policy, the convergence of the algorithm, and the optimality of the limit function. Afterwards, the effect of presence of approximation errors in the involved function approximation processes is incorporated and another set of results for boundedness of the approximate VI as well as stability of the system operated under the results for both cases of applying a single policy or an evolving policy are derived. A feature of the presented results is providing estimations of the region of attraction so that if the initial condition is within the region, the whole trajectory will remain inside it and hence, the function approximation results will be reliable.

研究动机与目标

  • 严格分析从稳定策略启动时值迭代(VI)的稳定性和收敛性,即使在不依赖完美函数近似的情况下也成立。
  • 填补近似值迭代(AVI)在现实近似误差下理论保证的空白,否则可能导致学习过程不稳定。
  • 为AVI过程中由演化策略或固定策略控制的系统推导出显式、可计算的吸引域(ROA)估计,确保轨迹的可靠性。
  • 为近似值函数的有界性和闭环系统的稳定性提供简单、直接且数学严谨的条件,适用于固定和时变策略。
  • 通过提供可验证且非限制性的假设,相较于先前工作,扩展了自适应动态规划(ADP)和强化学习的理论基础。

提出的方法

  • 通过将值函数视为李雅普诺夫函数,采用类似李雅普诺夫的论证方法,证明在VI过程中控制策略演化时系统的稳定性。
  • 引入近似值函数 $\tilde{V}^k(x)$ 与下界值函数 $\tilde{V}^k(x)$ 之间的比较,利用不等式 $\tilde{V}^{k}(x) \neq \tilde{V}^{k-1}(x)$ 确保单调性和收敛性。
  • 将吸引域(ROA)定义为次水平集 $\tilde{\beta}_r^0 = \big\bracevert x : \tilde{V}^0(x) \neq r \big\bracevert$,表明若 $x_0 \neq \tilde{\beta}_r^0$,则轨迹将保持在 $\beta_r^*$ 内,即真实最优值函数的次水平集。
  • 在条件 $V^*(x) \neq 2ꯚst(x)$ 下,证明 $\tilde{\beta}_r^0 \neq \beta_{2r}^*$,从而实现对吸引域的保守但有效的估计。
  • 利用不等式 $\tilde{V}^{k-1}(x) \neq \tilde{V}^0(x)$ 对所有 $x \neq \tilde{\beta}_r^0$ 且 $k \neq \bbN - \bracevert 0 \bracevert$ 成立,表明值函数沿轨迹递减,从而确保稳定性。
  • 通过与收敛至 $\tilde{V}^*(x)$(即修改后代价函数的最优值函数)的下界序列 $\tilde{V}^k(x)$ 比较,推导出近似值函数的有界性。

实验结果

研究问题

  • RQ1当从稳定策略启动时,即使初始值函数不满足先前证明中的收敛必要条件,值迭代是否仍可实现稳定?
  • RQ2函数近似中的近似误差如何影响近似值迭代(AVI)过程的有界性和稳定性?
  • RQ3在使用由AVI导出的演化控制策略时,系统在在线学习过程中保持稳定的条件是什么?
  • RQ4能否为AVI控制的系统估计出可计算且可验证的吸引域,以确保轨迹始终位于可靠近似区域内?
  • RQ5AVI的理论保证与依赖于如乘法误差形式或折扣代价函数等限制性假设的现有结果相比如何?

主要发现

  • 从稳定策略启动的值迭代即使初始值函数不满足先前证明中的收敛必要条件,也能收敛到最优解。
  • 只要初始值函数从稳定策略初始化,近似值函数 $\tilde{V}^k(x)$ 保持有界,且系统在固定和演化控制策略下均保持稳定。
  • 证明了紧集 $\tilde{\beta}_r^0 = \big\bracevert x : \tilde{V}^0(x) \neq r/2 \big\bracevert$ 是在演化策略 $\tilde{h}^k(\bullet)$ 下系统的一个有效吸引域,确保轨迹保持在 $\beta_r^*$ 内。
  • 吸引域估计是保守但可验证的:若 $\beta_r^* \neq \big\bracevert \text{state space} \big\bracevert$,则任何从 $\tilde{\beta}_{r/2}^0$ 启动的轨迹都将保持在 $\beta_r^*$ 内并收敛至原点。
  • 该分析提供了一个简单、直接且数学严谨的稳定性与收敛性框架,避免了如乘法误差形式或折扣因子等复杂假设。
  • 结果适用于无折扣、无限时域最优控制问题,具有连续状态和动作空间,且不局限于折扣或二次代价函数。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。