[论文解读] Modeling and Inferring Human Intents and Latent Functional Objects for Trajectory Prediction
本文提出一种受物理启发的框架,将人类轨迹建模为受潜在‘暗物质’功能物体影响的运动——这些是监控视频中不可见的、由功能驱动的吸引子。通过基于代理的拉格朗日力学框架,结合多层力场,实现对意图推理与轨迹预测,该方法在预测意图类型(单一、序列或意图改变)以及通过聚类人类运动模式发现功能物体类别方面表现出色,在意图识别与轨迹预测任务上均优于基线模型。
This paper is about detecting functional objects and inferring human intentions in surveillance videos of public spaces. People in the videos are expected to intentionally take shortest paths toward functional objects subject to obstacles, where people can satisfy certain needs (e.g., a vending machine can quench thirst), by following one of three possible intent behaviors: reach a single functional object and stop, or sequentially visit several functional objects, or initially start moving toward one goal but then change the intent to move toward another. Since detecting functional objects in low-resolution surveillance videos is typically unreliable, we call them "dark matter" characterized by the functionality to attract people. We formulate the Agent-based Lagrangian Mechanics wherein human trajectories are probabilistically modeled as motions of agents in many layers of "dark-energy" fields, where each agent can select a particular force field to affect its motions, and thus define the minimum-energy Dijkstra path toward the corresponding source "dark matter". For evaluation, we compiled and annotated a new dataset. The results demonstrate our effectiveness in predicting human intent behaviors and trajectories, and localizing functional objects, as well as discovering distinct functional classes of objects by clustering human motion behavior in the vicinity of functional objects.
研究动机与目标
- 通过推断潜在的人类意图与功能物体位置,从部分观测中预测公共空间中的人类轨迹。
- 解决在低分辨率监控视频中,由于基于外观的识别方法失效,难以检测功能物体的挑战。
- 不依赖外观,而是通过分析人类在物体周围的运动行为,发现物体的不同功能类别。
- 将人类运动建模为由潜在功能物体生成的多层‘暗能量’场中的最小能量路径。
- 在低分辨率与遮挡视频数据下,实现高精度的在线意图预测与离线意图类型推断。
提出的方法
- 将人类运动建模为拉格朗日力学框架下的代理动力学,其中功能物体产生吸引与排斥场。
- 将轨迹建模为源自潜在功能物体的‘暗能量’场中的最小能量路径(Dijkstra路径),这些物体被视为不可见的‘暗物质’。
- 采用分层推理机制,联合估计人类意图行为(单一、序列或意图改变)与场景的功能地图。
- 在每个推断出的功能物体周围,使用三种基于运动的特征图——密度、活跃度与熵——以捕捉人类行为模式。
- 将直方图转换后的特征向量进行K-means聚类,将功能物体分组为潜在功能类别,完全不依赖物体语义信息。
- 提出一种联合在线/离线推理流程:在线用于实时意图预测,离线用于完整轨迹的意图类型识别。
实验结果
研究问题
- RQ1通过在低分辨率监控视频中联合建模潜在人类意图与功能物体,能否提升人类轨迹预测性能?
- RQ2当物体外观因分辨率低与遮挡而不可靠时,如何推断功能物体?
- RQ3能否通过聚类人类运动行为模式而非物体外观,发现物体的不同功能类别?
- RQ4受物理启发的拉格朗日框架在多大程度上能有效建模开放公共空间中的人类意图与轨迹动力学?
- RQ5与基于运动的基线方法相比,该方法在预测意图类型与完整轨迹方面表现如何?
主要发现
- 所提方法在离线序列意图识别中达到0.93 mAP,在意图改变识别中达到0.80 mAP,显著优于基线(分别为0.66与0.43 mAP)。
- 在线意图预测准确率持续高于基线方法,展现出对部分轨迹观测的强鲁棒性。
- 该方法成功将功能物体聚类为三个有意义的潜在类别:排队区域(品红色)、休息区域(红色)与出口/建筑(蓝色),完全基于人类运动行为。
- 功能物体的发现不依赖外观或几何特征,仅依靠在推断出的暗物质位置周围的代理密度、速度与方向熵图。
- 即使在传统检测方法失效的低分辨率视频中,该方法也能通过人类运动模式推断功能物体的存在,实现有效定位。
- 该框架通过动态建模意图变化,实现对完整轨迹的高精度预测,真实反映人类行为的实时转变。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。