[论文解读] A numerical approach to stochastic reach-avoid problems for Markov Decision Processes.
该论文提出了一种数值方法,用于求解马尔可夫决策过程中高维随机可达-规避问题,通过将递归方程松弛为不等式,将值函数投影到高斯径向基函数上,并采样约束条件。该方法可实现对超长方体安全集和目标集的一步奖励的解析计算,显著优于基于网格的方法,并在非线性自主赛车路径规划问题上表现出有效性。
We consider finite horizon reach-avoid problems for discrete time stochastic systems with additive Gaussian mixture noise. Our goal is to approximate the optimal value function of such problems on dimensions higher than what can be handled via state-space gridding techniques. We achieve this by relaxing the recursive equations of the finite horizon reach-avoid optimal control problem into inequalities, projecting the optimal value function to a finite dimensional basis and sampling the associated infinite set of constraints. We focus on a specific value function parametrization using Gaussian radial basis functions that enables the analytical computation of the one-step reach-avoid reward in the case of hyper-rectangular safe and target sets, achieving significant computational benefits compared to state-space gridding. We analyze the performance of the overall method numerically by approximating simple reach-avoid control problems and comparing the results to benchmark controllers based on well-studied methods. The full power of the method is demonstrated on a nonlinear control problem inspired from constrained path planning for autonomous race cars.
研究动机与目标
- 解决在高维随机系统中有限时域可达-规避问题的挑战,其中状态空间网格化变得计算不可行。
- 为具有加性高斯混合噪声的系统开发一种可扩展的最优值函数近似方法。
- 通过使用高斯径向基函数的特定参数化,实现一步可达-规避奖励的解析计算。
- 在简单基准问题和复杂的非线性自主车辆控制问题上展示该方法的性能。
- 为随机最优控制提供一种计算高效的替代传统基于网格的方法。
提出的方法
- 将有限时域可达-规避最优控制问题的递归方程松弛为不等式,以实现数值近似。
- 将最优值函数投影到有限维的高斯径向基函数(RBF)基上,以实现可计算性。
- 通过采样由松弛不等式产生的无限约束集,形成有限的优化问题。
- 利用超长方体安全集和目标集的结构,解析计算一步可达-规避奖励,减少对数值积分的依赖。
- 将问题表述为可使用标准数值求解器求解的约束优化任务。
- 在低维问题上验证该方法,并将其扩展到高维非线性自主赛车路径规划问题。
实验结果
研究问题
- RQ1基于松弛的方法结合函数逼近是否能在高维随机可达-规避问题中优于传统的状态空间网格化?
- RQ2高斯径向基函数在具有超长方体集合的可达-规避设置中,能在多大程度上实现一步奖励的解析计算?
- RQ3与基准控制器相比,该方法在性能和计算效率方面表现如何?
- RQ4该方法能否有效扩展到非线性、高维系统,如自主车辆路径规划?
- RQ5约束采样和基函数选择对值函数近似精度和收敛性有何影响?
主要发现
- 通过高斯RBF参数化实现一步奖励的解析计算,该方法在计算上显著优于状态空间网格化。
- 数值结果表明,该方法在简单可达-规避问题上以高精度近似最优值函数,优于基准控制器。
- 该方法成功解决了自主赛车的复杂非线性路径规划问题,展示了其可扩展性和实际相关性。
- 使用高斯径向基函数可实现高维空间中值函数的高效且精确投影。
- 约束采样有效将无限维问题简化为有限维、可求解的优化问题,而无需牺牲解的质量。
- 对超长方体集合的一步奖励进行解析处理,消除了昂贵的数值积分,从而提高了计算效率。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。