Skip to main content
QUICK REVIEW

[论文解读] Approximating the Stationary Hamilton-Jacobi-Bellman Equation by Hierarchical Tensor Products

Mathias Oster, Leon Sallandt|arXiv (Cornell University)|Nov 1, 2019
Model Reduction and Neural Networks参考文献 2被引用 12
一句话总结

该论文提出了一种分层张量列车(TT)与多变量多项式逼近方法,用于求解由空间半离散化非线性抛物型PDE的无限时域最优控制问题所引出的高维稳态Hamilton-Jacobi-Bellman(HJB)方程。通过将策略迭代重新表述为线性双曲型PDE,并结合高维积分与Koopman算子,实现了对值函数的低秩逼近,且在黏性Burgers方程与Schlögl方程上验证了其收敛性与数值稳定性。

ABSTRACT

We treat infinite horizon optimal control problems by solving the associated stationary Hamilton-Jacobi-Bellman (HJB) equation numerically, for computing the value function and an optimal feedback area law. The dynamical systems under consideration are spatial discretizations of nonlinear parabolic partial differential equations (PDE), which means that the HJB is suffering from the curse of dimensions. To overcome numerical infeasability we use low-rank hierarchical tensor product approximation, or tree-based tensor formats, in particular tensor trains (TT tensors) and multi-polynomials, since the resulting value function is expected to be smooth. To this end we reformulate the Policy Iteration algorithm as a linearization of HJB equations. The resulting linear hyperbolic PDE remains the computational bottleneck due to high-dimensions. By the methods of characteristics it can be reformulated via the Koopman operator in the spirit of dynamic programming. We use a low rank tensor representation for approximation of the value function. The resulting operator equation is solved using high-dimensional quadrature, e.g. Variational Monte-Carlo methods. From the knowledge of the value function at computable samples $x_i$ we infer the function $ x \mapsto v (x)$. We investigate the convergence of this procedure. By controlling destabilized versions of viscous Burgers and Schloegl equations numerical evidences are given.

研究动机与目标

  • 解决由空间半离散化非线性抛物型PDE所控制的无限时域最优控制问题中的维数灾难问题。
  • 开发一种在高维下近似稳态HJB方程值函数的数值可行方法。
  • 利用低秩张量格式——特别是张量列车与多变量多项式——高效表示光滑的值函数。
  • 将策略迭代重新表述为线性双曲型PDE,并通过Koopman算子与高维积分求解。
  • 通过在不稳定黏性Burgers方程与Schlögl方程上的数值实验验证该方法。

提出的方法

  • 将稳态HJB方程重新表述为策略迭代过程,将非线性HJB方程线性化为一系列线性双曲型PDE。
  • 应用特征线法将线性双曲型PDE转化为基于Koopman算子的演化方程。
  • 使用分层张量格式(特别是张量列车TT与多变量多项式)表示值函数,以利用其低秩结构。
  • 采用高维积分技术(如变分蒙特卡洛)在选定点上计算值函数的采样点。
  • 从离散采样点中推断全局值函数 $ x \mapsto v(x) $,采用低秩张量逼近方法。
  • 通过在黏性Burgers方程与Schlögl方程的稳定与不稳定版本上应用该方法,控制解过程的数值稳定性。

实验结果

研究问题

  • RQ1低秩张量格式能否有效逼近由抛物型PDE引出的高维HJB方程的值函数?
  • RQ2通过Koopman算子重新表述策略迭代,是否能实现对高维下线性双曲型PDE的稳定且高效的求解?
  • RQ3结合TT逼近的高维积分方法在多大程度上保持了收敛性与精度?
  • RQ4该方法在具有不稳定控制的非线性PDE(如黏性Burgers方程与Schlögl方程)上的表现如何?
  • RQ5能否通过基于张量的插值方法,从稀疏且可计算的采样点中准确重构值函数?

主要发现

  • 所提出的方法成功利用低秩张量列车与多变量多项式逼近了HJB方程的值函数,克服了维数灾难问题。
  • 通过Koopman算子重新表述策略迭代,为线性双曲型PDE提供了稳定且计算上可行的求解路径。
  • 高维积分方法(如变分蒙特卡洛)能够在高维状态空间中实现值函数的精确采样。
  • 数值解算过程表现出收敛性,各测试案例中误差持续减少。
  • 在不稳定黏性Burgers方程与Schlögl方程上的数值实验表明,该方法在存在非线性与高维性的情况下仍保持稳定与高精度。
  • 通过基于张量的逼近方法,能从采样点中准确重构值函数,证实了低秩表示的有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。