[论文解读] Neural Foundations of Mental Simulation: Future Prediction of Latent Representations on Dynamic Scenes
本研究通过训练和评估感官-认知模型来预测动态、自然环境中的未来状态,探究了心理模拟的神经基础。结果表明,只有在预训练于多样化第一人称任务的视频基础模型潜在空间中预测未来状态的模型,才能在所有测试环境中准确匹配灵长类神经动力学和人类行为错误模式,表明心理模拟依赖于可重用的、动态的视觉表征,这些表征在跨环境的物理预测中得到优化。
Humans and animals have a rich and flexible understanding of the physical world, which enables them to infer the underlying dynamical trajectories of objects and events, plausible future states, and use that to plan and anticipate the consequences of actions. However, the neural mechanisms underlying these computations are unclear. We combine a goal-driven modeling approach with dense neurophysiological data and high-throughput human behavioral readouts to directly impinge on this question. Specifically, we construct and evaluate several classes of sensory-cognitive networks to predict the future state of rich, ethologically-relevant environments, ranging from self-supervised end-to-end models with pixel-wise or object-centric objectives, to models that future predict in the latent space of purely static image-based or dynamic video-based pretrained foundation models. We find strong differentiation across these model classes in their ability to predict neural and behavioral data both within and across diverse environments. In particular, we find that neural responses are currently best predicted by models trained to predict the future state of their environment in the latent space of pretrained foundation models optimized for dynamic scenes in a self-supervised manner. Notably, models that future predict in the latent space of video foundation models that are optimized to support a diverse range of sensorimotor tasks, reasonably match both human behavioral error patterns and neural dynamics across all environmental scenarios that we were able to test. Overall, these findings suggest that the neural mechanisms and behaviors of primate mental simulation are thus far most consistent with being optimized to future predict on dynamic, reusable visual representations that are useful for Embodied AI more generally.
研究动机与目标
- 识别支持动态环境灵活心理模拟的神经与行为系统中的归纳偏好。
- 评估最先进机器学习模型是否能复现灵长类前额叶皮层神经动力学与人类在预测任务中的行为模式。
- 确定哪种模型架构、损失函数和预训练方案最能预测心理模拟中被遮挡物理轨迹时的神经反应。
- 评估模型在多样化、高变化性环境以及训练分布之外新场景中的泛化能力。
- 探究在潜在空间中进行未来预测的模型是否能推断出输入中不可见的隐藏物理状态变量。
提出的方法
- 训练并评估多类感官-认知网络:端到端自监督模型(采用像素级或物体槽目标),以及在静态或动态基础模型潜在空间中进行未来预测的模型。
- 利用灵长类前额叶皮层在部分遮挡条件下的球体拦截任务(Mental-Pong)中获得的密集神经生理记录,作为模型预测的基准。
- 采用大规模人类行为数据(数千次比较)评估模型在未来预测与错误模式上的表现。
- 将模型预测与真实环境状态变量(包括视觉上隐藏的变量如速度)进行对比,以检验其隐含的物理理解能力。
- 在Physion基准的多样化动态场景上评估模型表现,涵盖刚体、软体(悬垂)及复杂交互场景。
- 使用前向动力学头实现多时间步的预测滚动,以评估长时程泛化能力。
实验结果
研究问题
- RQ1哪种模型架构与训练目标最能预测灵长类前额叶皮层在被遮挡物理轨迹心理模拟过程中的神经动力学?
- RQ2在潜在空间中进行未来预测训练的模型能否推断出输入中不可见的隐藏物理状态变量(如速度)?
- RQ3模型在超出其训练分布的新环境与高变化性场景中如何泛化?
- RQ4在第一人称视频数据上预训练的基础模型是否优于在静态图像或像素级重建任务上预训练的模型,以预测神经与行为反应?
- RQ5可重用、解耦且以物体为中心的表征在实现准确且生物合理心理模拟中起到何种作用?
主要发现
- 只有在预训练于多样化第一人称感官运动任务的视频基础模型潜在空间中进行未来预测的模型,才能在所有测试环境中准确匹配灵长类神经动力学与人类行为错误模式。
- 在像素级重建或物体槽预测上训练的模型即使扩大规模,也难以实现良好泛化,表明规模本身不足以支持心理模拟。
- 表现最佳的模型即使未显式训练预测隐藏物理状态变量(如速度),也能成功预测,表明其具备隐含的物理理解能力。
- 基础模型潜在空间并非等同:在动态第一人称视频数据(如Ego4D)上预训练的模型,优于在静态图像或非第一人称数据上预训练的模型。
- 与神经数据匹配度最高的模型类别,也是唯一在软体‘Drape’场景中表现良好的模型,而其他所有模型在此场景中均显著受挫。
- 神经反应最能被那些在可重用、动态潜在空间中预测未来状态的模型所预测,表明心理模拟的核心归纳偏好依赖于此类表征。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。