[论文解读] Active Inference and Intentional Behaviour
本文提出,通过归纳规划实现的主动推理可使意图性行为——即由潜在状态空间目标引导的目标导向行为——自然涌现,从而将其与反应性行为和感知性行为区分开来。通过体外神经元培养和网格世界基准的模拟,研究显示:在潜在状态空间中直接指定目标,可实现通过自由能最小化进行的快速、高效且长时程的规划,其性能优于传统的基于奖励的强化学习方法。
Recent advances in theoretical biology suggest that basal cognition and sentient behaviour are emergent properties of in vitro cell cultures and neuronal networks, respectively. Such neuronal networks spontaneously learn structured behaviours in the absence of reward or reinforcement. In this paper, we characterise this kind of self-organisation through the lens of the free energy principle, i.e., as self-evidencing. We do this by first discussing the definitions of reactive and sentient behaviour in the setting of active inference, which describes the behaviour of agents that model the consequences of their actions. We then introduce a formal account of intentional behaviour, that describes agents as driven by a preferred endpoint or goal in latent state-spaces. We then investigate these forms of (reactive, sentient, and intentional) behaviour using simulations. First, we simulate the aforementioned in vitro experiments, in which neuronal cultures spontaneously learn to play Pong, by implementing nested, free energy minimising processes. The simulations are then used to deconstruct the ensuing predictive behaviour, leading to the distinction between merely reactive, sentient, and intentional behaviour, with the latter formalised in terms of inductive planning. This distinction is further studied using simple machine learning benchmarks (navigation in a grid world and the Tower of Hanoi problem), that show how quickly and efficiently adaptive behaviour emerges under an inductive form of active inference.
研究动机与目标
- 在主动推理框架中,通过在潜在状态空间中定义目标而非使用奖励函数,形式化意图性行为。
- 通过自由能最小化和预测编码,区分反应性、感知性和意图性行为。
- 证明归纳规划可实现在神经元模拟和经典规划任务中的快速、高效且目标导向的行为。
- 验证体外神经元培养中的自证行为源于自由能最小化,且无需外部奖励。
- 表明归纳规划结合了长期规划的覆盖范围与短期推理的效率。
提出的方法
- 将智能体建模为通过最小化变分自由能来解释神经网络中自组织与学习的机制。
- 通过嵌套的自由能最小化过程,模拟体外神经元培养在类似乒乓球游戏任务中的行为。
- 通过归纳规划定义意图性行为,即在潜在状态空间中将目标指定为期望的终点。
- 使用逆向归纳法计算最小化期望自由能以达到目标状态的策略。
- 将归纳规划应用于网格世界导航和河内塔问题等基准任务。
- 通过将目标状态视为未来潜在状态的精确先验,将规划形式化为推理。
实验结果
研究问题
- RQ1在主动推理框架中,如何形式化区分意图性行为与反应性和感知性行为?
- RQ2归纳规划在实现高效、长时程目标导向行为中发挥何种作用?
- RQ3体外神经元培养是否可在无强化学习或外部奖励的情况下表现出适应性、目标导向行为?
- RQ4与基于奖励的方法相比,通过在潜在状态空间中指定目标如何提升规划效率?
- RQ5自由能最小化在多大程度上可解释生物和人工系统中的自证行为与目的论行为?
主要发现
- 体外神经元培养通过自由能最小化自发学会玩乒乓球游戏,且无需外部奖励或强化学习。
- 反应性、感知性和意图性行为之间的区分,自然地从自由能最小化的结构及归纳规划的使用中浮现。
- 归纳规划在河内塔和网格世界任务中实现了快速收敛至最优解,优于标准的无模型和有模型强化学习方法。
- 意图性行为被形式化为通过精确先验最小化期望自由能,从而实现长时程规划。
- 该框架解释了智能体如何仅通过最小化意外性与概率推理,实现复杂的目标导向行为。
- 结果支持如下假设:基础认知与感知性源于神经系统的自由能最小化,即使没有显式奖励信号。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。