Skip to main content
QUICK REVIEW

[论文解读] Inverse Optimal Control with Incomplete Observations.

Wanxin Jin, Dana Kulić|arXiv (Cornell University)|Mar 21, 2018
Robotics and Sensor-Based Localization参考文献 23被引用 14
一句话总结

本文提出了一种新颖的逆最优控制方法,通过引入一个恢复矩阵,将观测到的状态与代价函数特征权重关联起来,从而从不完整的轨迹观测中恢复代价函数权重。该方法在仅需最少观测数据的情况下,实现了稳定、准确且鲁棒的代价函数学习,在模拟机器人机械臂任务中验证了其性能优于当前最先进方法。

ABSTRACT

In this article, we consider the inverse optimal control problem given incomplete observations of an optimal trajectory. We hypothesize that the cost function is constructed as a weighted sum of relevant features (or basis functions). We handle the problem by proposing the recovery matrix, which establishes a relationship between available observations of the trajectory and weights of given candidate features. The rank of the recovery matrix indicates whether a subset of relevant features can be found among the candidate features and the corresponding weights can be recovered. Additional observations tend to increase the rank of the recovery matrix, thus enabling cost function recovery. We also show that the recovery matrix can be computed iteratively. Based on the recovery matrix, a methodology for using incomplete observations of the trajectory to recover the weights of specified features is established, and an efficient algorithm for recovering the feature weights by finding the minimal required observations is developed. We apply the proposed algorithm to learning the cost function of a simulated robot manipulator conducting free-space motions. The results demonstrate the stable, accurate and robust performance of the proposed approach compared to state of the art techniques.

研究动机与目标

  • 解决仅可观测部分轨迹时逆最优控制的挑战。
  • 识别在不完整观测下可从候选特征中可靠恢复的子集。
  • 开发一种高效算法,以最小化实现准确权重恢复所需的观测数量。
  • 建立在观测不确定性下进行代价函数学习的系统性框架。

提出的方法

  • 该方法引入一个恢复矩阵,将代价函数中基函数加权和的观测轨迹状态映射到特征权重。
  • 恢复矩阵的秩决定了是否能够唯一识别出相关特征子集并恢复其权重。
  • 恢复矩阵通过迭代计算,支持在新观测可用时实现增量式学习。
  • 通过分析矩阵秩,该算法识别出实现完整权重恢复所需的最小观测集合。
  • 该方法假设代价函数在特征权重上线性,从而可通过线性代数实现高效优化。
  • 该方法通过矩阵秩分析实现特征相关性识别,确保仅选择具有信息量的特征进行恢复。

实验结果

研究问题

  • RQ1能否从最优轨迹的不完整观测中准确恢复代价函数?
  • RQ2在不完整观测条件下,哪些候选特征子集可以被唯一识别并赋予权重?
  • RQ3如何确定实现稳定且准确权重恢复所需的最小观测数量?
  • RQ4恢复矩阵在实现增量式与鲁棒逆最优控制中起到什么作用?

主要发现

  • 恢复矩阵的秩可作为可靠指标,判断是否能从不完整观测中唯一恢复特征子集及其权重。
  • 额外的观测可提高矩阵秩,从而增强完整代价函数恢复的可能性。
  • 恢复矩阵的迭代计算支持高效、增量式学习,且数据需求极低。
  • 所提出的算法在模拟机器人机械臂任务中实现了稳定、准确且鲁棒的特征权重恢复性能。
  • 在不完整观测条件下,该方法在准确性和鲁棒性方面优于当前最先进技术。
  • 该方法成功识别出恢复代价函数权重所需的最小观测集合,显著降低了数据依赖性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。