[论文解读] Modeling Human Understanding of Complex Intentional Action with a Bayesian Nonparametric Subgoal Model
本文提出了一种贝叶斯非参数子目标模型,通过假设对未知子目标序列的近似理性规划,从观察到的复杂动作中推断分层动作结构。该模型准确预测了人类的子目标推断,在行为实验和人工用户辅助任务中显著优于其他模型,实现了更快速、更稳定且更准确的支持,且所需观测数据更少。
Most human behaviors consist of multiple parts, steps, or subtasks. These structures guide our action planning and execution, but when we observe others, the latent structure of their actions is typically unobservable, and must be inferred in order to learn new skills by demonstration, or to assist others in completing their tasks. For example, an assistant who has learned the subgoal structure of a colleague's task can more rapidly recognize and support their actions as they unfold. Here we model how humans infer subgoals from observations of complex action sequences using a nonparametric Bayesian model, which assumes that observed actions are generated by approximately rational planning over unknown subgoal sequences. We test this model with a behavioral experiment in which humans observed different series of goal-directed actions, and inferred both the number and composition of the subgoal sequences associated with each goal. The Bayesian model predicts human subgoal inferences with high accuracy, and significantly better than several alternative models and straightforward heuristics. Motivated by this result, we simulate how learning and inference of subgoals can improve performance in an artificial user assistance task. The Bayesian model learns the correct subgoals from fewer observations, and better assists users by more rapidly and accurately inferring the goal of their actions than alternative approaches.
研究动机与目标
- 建模人类如何从一系列有意图动作中推断潜在的子目标结构。
- 开发一种计算框架,利用非参数贝叶斯方法捕捉分层规划与理性动作推断。
- 检验该模型是否比其他模型和启发式方法更准确地预测人类的子目标推断。
- 评估该模型在人工用户辅助任务中的实用性,其中模型通过学习子目标来实时支持用户。
- 证明子目标推断可实现更快速、更稳定的辅助,且训练数据极少。
提出的方法
- 该模型使用分层马尔可夫决策过程(MDP),其中动作是通过对未知子目标序列进行近似理性规划而生成的。
- 子目标序列使用狄利克雷过程先验进行建模,从而实现在子目标数量和组成上的非参数推断。
- 采用自定义的马尔可夫链蒙特卡洛(MCMC)方法,从观测到的动作序列中联合推断子目标序列的数量和结构。
- 该模型通过贝叶斯推断计算对目标和子目标序列的后验概率,结合逆向规划引入理性约束。
- 在用户辅助任务中,模型通过少量示范学习子目标结构,并利用其预测和辅助正在进行的部分动作序列。
- 性能通过协作任务进行评估,助手需推断子目标并协助用户,奖励机制对动作成本进行惩罚。
实验结果
研究问题
- RQ1人类如何从复杂动作序列中推断子目标的数量和组成?
- RQ2具有分层规划的贝叶斯非参数模型是否能比其他模型更准确地预测人类的子目标推断?
- RQ3子目标推断在多大程度上提升了人工用户辅助任务中的性能?
- RQ4模型性能如何随训练示范数量的变化而变化?
- RQ5即使数据有限,该模型的子目标推断是否依然稳定可靠?
主要发现
- 贝叶斯非参数子目标模型以高精度预测了人类的子目标推断,显著优于其他模型和启发式方法。
- 仅使用两个或更多输入序列,该模型在用户辅助任务中便达到了接近真实水平的性能,展现出强大的数据效率。
- 该模型在重复试验中表现稳定,方差极低,而其他模型在不同输入序列下表现出高度不稳定性。
- 即使训练数据极少,该模型仍使助手比缺乏子目标知识或采用固定子目标结构的模型更快速、更准确地协助用户。
- 该模型推断子目标边界和分层结构的能力,带来了更好的泛化能力和上下文敏感的动作理解。
- 该模型的非参数特性使其无需预先指定即可学习正确的子目标数量及其序列,从而提升了灵活性和鲁棒性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。