[论文解读] The relationship between dynamic programming and active inference: the discrete, finite-horizon case.
本文证明,在离散、有限时域的部分可观察马尔可夫决策过程(POMDP)中,贝尔曼方程下的动态规划是主动推理的极限情况。通过最小化期望自由能,主动推理统一了奖励最大化与模糊性减少,揭示了探索与利用行为如何自然地从单一变分推理框架中产生。
Active inference is a normative framework for generating behaviour based upon the free energy principle, a theory of self-organisation. This framework has been successfully used to solve reinforcement learning and stochastic control problems, yet, the formal relation between active inference and reward maximisation has not been fully explicated. In this paper, we consider the relation between active inference and dynamic programming under the Bellman equation, which underlies many approaches to reinforcement learning and control. We show that, on partially observable Markov decision processes, dynamic programming is a limiting case of active inference. In active inference, agents select actions to minimise expected free energy. In the absence of ambiguity about states, this reduces to matching expected states with a target distribution encoding the agent's preferences. When target states correspond to rewarding states, this maximises expected reward, as in reinforcement learning. When states are ambiguous, active inference agents will choose actions that simultaneously minimise ambiguity. This allows active inference agents to supplement their reward maximising (or exploitative) behaviour with novelty-seeking (or exploratory) behaviour. This clarifies the connection between active inference and reinforcement learning, and how both frameworks may benefit from each other.
研究动机与目标
- 阐明主动推理与动态规划在序列决策中的形式关系。
- 研究当状态模糊性缺失时,主动推理如何退化为动态规划。
- 展示主动推理如何通过自由能最小化自然地整合奖励最大化与探索行为。
- 在统一的变分推理框架下,将强化学习与主动推理统一起来。
提出的方法
- 将主动推理形式化为一种最小化信念状态上期望自由能的变分推理过程。
- 应用贝尔曼方程来建模部分可观察马尔可夫决策过程(POMDP)中的价值函数。
- 推导出最小化期望自由能退化为最小化期望成本的条件,即等价于动态规划。
- 表明当状态不确定性为零时,自由能最小化与标准强化学习中的奖励最大化一致。
- 证明在存在模糊性时,智能体同时最小化期望成本与模型不确定性,从而实现内在探索。
- 利用自由能分解表明,主动推理通过包含认识论不确定性,将动态规划推广至更一般情形。
实验结果
研究问题
- RQ1在有限时域POMDP中,主动推理与贝尔曼方程下的动态规划有何关系?
- RQ2在何种条件下,主动推理会退化为动态规划?
- RQ3主动推理如何在一个统一的原理性框架内整合奖励最大化与探索行为?
- RQ4认识论不确定性在主动推理中如何影响行为的形成?
主要发现
- 当状态不确定性可忽略时,主动推理退化为动态规划,即智能体通过最小化期望成本来行动。
- 在无模糊性的情况下,最小化期望自由能等价于最大化期望奖励,与标准强化学习一致。
- 当状态存在模糊性时,主动推理智能体同时最小化期望成本与模型不确定性,从而实现内在探索。
- 该框架自然地平衡了利用(奖励最大化)与探索(模糊性减少),统一了强化学习中的两种核心行为。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。