Skip to main content
QUICK REVIEW

[论文解读] Inverse Optimal Control Adapted to the Noise Characteristics of the Human Sensorimotor System

M. Schultheis, Dominik Straub|arXiv (Cornell University)|Oct 21, 2021
Gene Regulatory Network Analysis被引用 7
一句话总结

本文提出了一种新颖的逆最优控制框架,通过考虑信号依赖性噪声(即运动变异性随控制信号幅值而变化)来推断人类感觉运动行为背后的代价函数。利用概率POMDP建模与矩匹配近似方法,该方法即使在部分可观测条件下,也能准确从观测轨迹中恢复代价参数,从而调和了规范性与描述性的人体运动控制模型。

ABSTRACT

Computational level explanations based on optimal feedback control with signal-dependent noise have been able to account for a vast array of phenomena in human sensorimotor behavior. However, commonly a cost function needs to be assumed for a task and the optimality of human behavior is evaluated by comparing observed and predicted trajectories. Here, we introduce inverse optimal control with signal-dependent noise, which allows inferring the cost function from observed behavior. To do so, we formalize the problem as a partially observable Markov decision process and distinguish between the agent's and the experimenter's inference problems. Specifically, we derive a probabilistic formulation of the evolution of states and belief states and an approximation to the propagation equation in the linear-quadratic Gaussian problem with signal-dependent noise. We extend the model to the case of partial observability of state variables from the point of view of the experimenter. We show the feasibility of the approach through validation on synthetic data and application to experimental data. Our approach enables recovering the costs and benefits implicit in human sequential sensorimotor behavior, thereby reconciling normative and descriptive approaches in a computational framework.

研究动机与目标

  • 为解决在噪声具有信号依赖性时,缺乏从人类感觉运动控制中推断代价函数的方法这一问题。
  • 形式化描述在部分可观测设置下,智能体内部推断与实验者推断问题之间的区别。
  • 开发一种基于似然的、可处理的逆最优控制方法,适用于信号依赖性噪声,超越标准LQG假设的限制。
  • 在合成数据与实验数据上验证该方法,证明其在状态变量部分可观测条件下的鲁棒性。
  • 调和计算运动控制中规范性(最优控制)与描述性(行为数据)方法之间的差异。

提出的方法

  • 将正向问题形式化为具有信号依赖性噪声的局部可观测马尔可夫决策过程(POMDP),同时建模智能体与实验者的推断过程。
  • 推导出包含信号依赖性噪声的概率状态演化模型,通过矩匹配近似非高斯不确定性。
  • 引入对称KL散度度量以评估观测轨迹与模拟轨迹之间的相似性,替代RMSE以提升鲁棒性。
  • 应用最大似然估计(MLE)从观测运动轨迹中推断代价函数参数。
  • 使用蒙特卡洛滚动仿真,验证矩匹配近似与经验轨迹分布的一致性。
  • 将框架扩展至部分可观测状态设置,将速度与加速度视为隐变量,并在不确定性条件下评估参数恢复效果。

实验结果

研究问题

  • RQ1逆最优控制能否适配于信号依赖性噪声,这是人类运动变异性的一个关键特征?
  • RQ2当状态变量仅部分可观测时,能否从观测轨迹中准确恢复代价函数参数?
  • RQ3在存在信号依赖性噪声的情况下,矩匹配近似对轨迹分布估计的影响如何?
  • RQ4在部分可观测条件下,参数估计误差如何影响模拟轨迹与观测轨迹之间的相似性?
  • RQ5该方法能否在不知晓代价函数先验知识的前提下,可靠地推断人类序列感觉运动行为背后的代价?

主要发现

  • 该方法在部分可观测条件下,从合成数据中成功恢复了代价函数参数,均方根误差(RMSE)极低,中位数RMSE分别为:奖励(r)4.0×10⁻²,速度惩罚(v)1.1×10⁻¹,加速度惩罚(f)2.6×10⁻¹。
  • 即使参数估计不够精确,模拟轨迹与观测轨迹之间的对称KL散度仍保持较低水平(约10⁻³),表明轨迹层面的相似性具有鲁棒性。
  • 矩匹配近似与经验轨迹分布高度一致(对称KL差异:1.60×10⁻³),显著优于采用加法噪声的基线方法(KL差异:6.05)。
  • 在部分可观测条件下,参数估计精度下降,尤其在速度与加速度惩罚方面,这是由于轨迹存在歧义所致,但对轨迹相似性影响不大。
  • 该方法在1000组随机参数配置下泛化能力良好,中位数MLE估计值与真实值高度一致,表明具有强大的统计一致性。
  • 该框架能够可靠地推断人类运动行为中隐含的代价与收益,弥合了感觉运动控制中规范性与描述性模型之间的鸿沟。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。