[论文解读] Risk sensitive nonlinear optimal control with measurement uncertainty.
本文提出了一种风险敏感的非线性最优控制框架,通过优化结合期望成本与高阶矩的性能准则,显式地考虑了过程噪声和测量噪声。该方法推导出一种依赖于估计误差和过程噪声协方差的仿射反馈控制律,突破了确定性等价原理,实现了在不确定环境下的鲁棒控制,已在两自由度机械臂的路径点任务和接触任务中得到验证。
We present an algorithm to synthesize locally-optimal feedback controllers that take into account additive process and measurement noise for nonlinear stochastic optimal control problems. The algorithm is based on an exponential performance criteria which allows to optimize not only the expected value of the cost, but also a linear combination of its higher order moments; thereby, the cost of uncertainty can be taken into account for synthesis of robust or risk-sensitive policies. The method constructs an affine feedback control law, whose gains explicitly depend upon the covariance of the estimation errors and process noise. Despite the fact that controller and observer are designed separately, the measurement noise variance enters the optimal control, therefore generating feedback laws that do not rely on the Certainty Equivalence Principle. The capabilities of the approach are illustrated in simulation on a two degree of freedom (DOF) manipulator, first in a waypoint task and then in a task where the manipulator goes in contact with its environment.
研究动机与目标
- 开发一种反馈控制综合方法,以在非线性随机最优控制问题中同时考虑过程噪声和测量噪声。
- 超越仅最小化期望成本,通过引入成本函数的高阶矩,实现风险敏感控制。
- 设计一种不依赖于确定性等价原理的控制器,即使控制器与观测器分别设计亦成立。
- 在涉及不确定性的复杂任务中,如与环境的接触,验证该方法的有效性。
- 提供一种计算上可行的方法,用于在不确定性下合成局部最优的反馈控制律。
提出的方法
- 该方法采用指数型性能准则,将期望成本与成本的高阶矩的风险敏感惩罚相结合。
- 其形式化了一种仿射反馈控制律,其增益显式依赖于估计误差和过程噪声的协方差。
- 通过非线性系统动力学的局部近似推导控制器,实现可计算的优化。
- 控制器与观测器的设计相互分离,但测量噪声方差仍被嵌入最优控制律中。
- 该方法使用一种类似Riccati的方程来计算考虑不确定性传播的反馈增益。
- 该算法在仿真中应用于两自由度机械臂,任务包括路径点跟踪和环境接触。
实验结果
研究问题
- RQ1如何通过考虑成本函数的高阶矩,将非线性最优控制扩展至风险敏感控制?
- RQ2在不依赖确定性等价原理的前提下,反馈控制性能在过程噪声和测量噪声下能提升到何种程度?
- RQ3即使观测器与控制器独立设计,控制器是否仍能有效利用测量噪声统计特性以增强鲁棒性?
- RQ4该方法在涉及动态不确定性和环境交互的任务中表现如何?
- RQ5在性能准则中引入方差和更高阶矩,对控制策略合成有何影响?
主要发现
- 所提出的算法成功生成了依赖于过程噪声和估计误差协方差的反馈控制律,实现了风险敏感控制。
- 通过引入成本的高阶矩,该方法在不确定性下提供了比仅考虑期望成本更全面的性能度量。
- 由于控制律中显式包含了测量噪声,即使观测器与控制器分别设计,控制器也不依赖于确定性等价原理。
- 在两自由度机械臂上的仿真结果表明,该方法在路径点跟踪和环境接触任务中均实现了稳定且鲁棒的性能。
- 通过在反馈策略中显式考虑估计成本和过程噪声,该方法显著提升了不确定性下的控制性能。
- 仿射反馈结构支持高效计算和局部最优性,使该方法适用于不确定环境中实时应用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。