Skip to main content
QUICK REVIEW

[论文解读] A neural network based policy iteration algorithm with global $H^2$-superlinear convergence for stochastic games on domains

Kazufumi Ito, Christoph Reisinger|arXiv (Cornell University)|Jun 5, 2019
Model Reduction and Neural Networks参考文献 45被引用 14
一句话总结

该论文提出了一种基于神经网络的策略迭代算法,用于在复杂域上求解高维随机HJBI方程,通过将该格式解释为一种非精确牛顿法,实现了全局H²-超线性收敛。该方法结合神经网络逼近与策略迭代来求解半线性PDE,确保了值函数和反馈控制均以最优速率收敛。

ABSTRACT

In this work, we propose a class of numerical schemes for solving semilinear Hamilton-Jacobi-Bellman-Isaacs (HJBI) boundary value problems which arise naturally from exit time problems of diffusion processes with controlled drift. We exploit policy iteration to reduce the semilinear problem into a sequence of linear Dirichlet problems, which are subsequently approximated by a multilayer feedforward neural network ansatz. We establish that the numerical solutions converge globally in the $H^2$-norm, and further demonstrate that this convergence is superlinear, by interpreting the algorithm as an inexact Newton iteration for the HJBI equation. Moreover, we construct the optimal feedback controls from the numerical value functions and deduce convergence. The numerical schemes and convergence results are then extended to HJBI boundary value problems corresponding to controlled diffusion processes with oblique boundary reflection. Numerical experiments on the stochastic Zermelo navigation problem are presented to illustrate the theoretical results and to demonstrate the effectiveness of the method.

研究动机与目标

  • 开发一种无网格的数值格式,用于求解随机控制问题中出现的半线性Hamilton-Jacobi-Bellman-Isaacs(HJBI)方程。
  • 克服传统有限差分法或有限元法在维度灾难和几何复杂性方面的局限性。
  • 为基于神经网络的策略迭代算法建立全局H²-收敛性及超线性收敛速率。
  • 将该方法扩展至处理斜导数边界条件,并从数值解中构造最优反馈控制。
  • 通过将该算法解释为非精确半光滑牛顿法,为收敛速率提供理论依据。

提出的方法

  • 该方法使用策略迭代将半线性HJBI问题转化为一系列线性狄利克雷问题或斜导数问题。
  • 每个线性子问题均采用多层前馈神经网络作为试函数空间,实现无网格逼近。
  • 该算法被解释为HJBI算子的非精确牛顿迭代,包含广义微分与拟牛顿更新。
  • 构建了一个扰动线性系统以考虑求解的不精确性,误差界由趋于零的序列ηk+1控制。
  • 在H²(Ω)范数下分析收敛性,利用椭圆斜导数问题的正则性理论。
  • 通过神经网络对值函数逼近的梯度推导出最优反馈控制。

实验结果

研究问题

  • RQ1基于神经网络的策略迭代格式能否在具有复杂几何形状的域上实现HJBI方程的全局H²-收敛?
  • RQ2所提出的算法是否在H²-范数下表现出超线性收敛?若然,其成立条件为何?
  • RQ3该方法能否扩展至处理斜导数边界条件,同时保持收敛速率?
  • RQ4求解线性子问题中的不精确性如何影响策略迭代循环的全局收敛性与收敛速率?
  • RQ5能否从值函数的神经网络逼近中可靠地构造出最优反馈控制?

主要发现

  • 数值解在H²(Ω)范数下全局收敛于HJBI方程的真实解。
  • 收敛速率为q-超线性,意味着随着迭代进行,误差的下降速度快于任意线性速率。
  • 通过将策略迭代解释为非精确半光滑牛顿法,该算法实现了H²-超线性收敛。
  • 在满足适当正则性条件的前提下,该方法在斜导数边界条件下仍保持收敛性。
  • 从神经网络值函数梯度导出的反馈控制收敛于最优控制。
  • 线性求解中的误差由趋于零的序列ηk+1控制,确保了非精确牛顿步在极限下仍具有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。