Skip to main content
QUICK REVIEW

[论文解读] Meshless discretization of LQ-type stochastic control problems

Ralf Banisch, Carsten Hartmann|arXiv (Cornell University)|Sep 28, 2013
Markov Chains and Monte Carlo Methods参考文献 39被引用 3
一句话总结

本论文提出了一种针对不定时域上线性二次型(LQ)类随机控制问题的无网格Galerkin离散化方法,通过对数变换将问题转化为线性两点边值问题。该方法实现了强L²误差界,并与马尔可夫决策问题对偶,适用于分子动力学等高维系统,采用共性函数基进行求解。

ABSTRACT

We propose a novel Galerkin discretization scheme for stochastic optimal control problems on an indefinite time horizon. The control problems are linear-quadratic in the controls, but possibly nonlinear in the state variables, and the discretization is based on the fact that problems of this kind can be transformed into linear boundary value problems by a logarithmic transformation. We show that the discretized linear problem is dual to a Markov decision problem, the precise form of which depends on the chosen Galerkin basis. We prove a strong error bound in $L^{2}$ for the general scheme and discuss two special cases: a variant of the known Markov chain approximation obtained from a basis of characteristic functions of a box discretization, and a sparse approximation that uses the basis of committor functions of metastable sets of the dynamics; the latter is particularly suited for high-dimensional systems, e.g., control problems in molecular dynamics. We illustrate the method with several numerical examples, one being the optimal control of Alanine dipeptide to its helical conformation.

研究动机与目标

  • 开发一种无网格离散化方法,用于LQ型随机控制问题,避免使用传统的基于网格的PDE求解器。
  • 解决PDE方法在高维系统中计算效率低下的问题。
  • 通过基函数的选择,建立Galerkin离散化与马尔可夫决策问题之间的对偶性。
  • 为通用方法提供强L²误差界,确保收敛性。
  • 通过稀疏基(如亚稳态集的共性函数)实现高维控制问题的高效求解,特别适用于分子动力学系统。

提出的方法

  • 该方法采用对数变换,将原始的非线性HJB方程转化为线性两点边值问题。
  • 通过用户定义的函数基实施Galerkin投影,将连续问题转化为离散线性系统。
  • 证明所得离散问题与马尔可夫决策问题对偶,其转移率由基函数和系统动力学显式导出。
  • 分析了两种特定基的选择:用于盒离散化的特征函数(得到马尔可夫链近似),以及用于稀疏、亚稳态感知离散化的共性函数。
  • 通过最佳逼近误差控制,确保强L²误差界,证明见附录B。
  • 通过里程碑法和路径概率公式,实现对离散化运行成本的采样,详见附录E。

实验结果

研究问题

  • RQ1无网格Galerkin方法能否有效应用于具有非线性状态依赖性的LQ型随机控制问题?
  • RQ2Galerkin基的选择如何影响对偶的马尔可夫决策问题及最终离散化的精度?
  • RQ3该方法的收敛行为如何?能否建立严格的L²误差界?
  • RQ4该方法能否高效处理高维系统,如分子动力学中的系统?
  • RQ5基于共性函数的基与标准马尔可夫链近似相比,在稀疏性和精度方面表现如何?

主要发现

  • 所提出的Galerkin格式实现了强L²误差界,通过最佳逼近误差控制证明,确保在适当基选择下收敛。
  • 离散化问题与马尔可夫决策问题对偶,其转移率由Galerkin基和系统参数显式导出。
  • 使用盒离散化下的特征函数可恢复已知的马尔可夫链近似,验证了方法的一致性。
  • 利用亚稳态集的共性函数可实现稀疏的高维离散化,特别适用于具有明显亚稳态结构的系统。
  • 在一维三阱势和丙氨酸二肽上的数值结果表明,该方法能高精度计算构象转变的最优控制力。
  • 该方法可通过里程碑法高效采样离散化运行成本,如公式(34)所形式化。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。