[Paper Review] The relationship between dynamic programming and active inference: the discrete, finite-horizon case.
This paper establishes that dynamic programming under the Bellman equation is a limiting case of active inference in discrete, finite-horizon partially observable Markov decision processes. By minimizing expected free energy, active inference unifies reward maximization and ambiguity reduction, revealing how exploratory and exploitative behaviors emerge naturally from a single variational inference framework.
Active inference is a normative framework for generating behaviour based upon the free energy principle, a theory of self-organisation. This framework has been successfully used to solve reinforcement learning and stochastic control problems, yet, the formal relation between active inference and reward maximisation has not been fully explicated. In this paper, we consider the relation between active inference and dynamic programming under the Bellman equation, which underlies many approaches to reinforcement learning and control. We show that, on partially observable Markov decision processes, dynamic programming is a limiting case of active inference. In active inference, agents select actions to minimise expected free energy. In the absence of ambiguity about states, this reduces to matching expected states with a target distribution encoding the agent's preferences. When target states correspond to rewarding states, this maximises expected reward, as in reinforcement learning. When states are ambiguous, active inference agents will choose actions that simultaneously minimise ambiguity. This allows active inference agents to supplement their reward maximising (or exploitative) behaviour with novelty-seeking (or exploratory) behaviour. This clarifies the connection between active inference and reinforcement learning, and how both frameworks may benefit from each other.
Motivation & Objective
- To clarify the formal relationship between active inference and dynamic programming in sequential decision-making.
- To investigate how active inference reduces to dynamic programming when state ambiguity is absent.
- To demonstrate how active inference naturally incorporates both reward maximization and exploration through free energy minimization.
- To unify reinforcement learning and active inference under a common variational inference framework.
Proposed method
- Formalizing active inference as a variational inference process minimizing expected free energy over belief states.
- Applying the Bellman equation to model value functions in partially observable Markov decision processes (POMDPs).
- Deriving conditions under which minimizing expected free energy reduces to minimizing expected cost, equivalent to dynamic programming.
- Showing that when state uncertainty is zero, free energy minimization aligns with reward maximization in standard RL.
- Demonstrating that in the presence of ambiguity, agents minimize both expected cost and model uncertainty, enabling intrinsic exploration.
- Using the free energy decomposition to show that active inference generalizes dynamic programming by including epistemic uncertainty.
Experimental results
Research questions
- RQ1How does active inference relate to dynamic programming under the Bellman equation in finite-horizon POMDPs?
- RQ2Under what conditions does active inference reduce to dynamic programming?
- RQ3How does active inference integrate reward maximization and exploration within a single principled framework?
- RQ4What role does epistemic uncertainty play in shaping behavior within active inference?
Key findings
- Active inference reduces to dynamic programming when state uncertainty is negligible, meaning agents act to minimize expected cost.
- In the absence of ambiguity, minimizing expected free energy is equivalent to maximizing expected reward, aligning with standard reinforcement learning.
- When states are ambiguous, active inference agents minimize both expected cost and model uncertainty, leading to intrinsic exploration.
- The framework naturally balances exploitation (reward maximization) and exploration (ambiguity reduction), unifying two core RL behaviors.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.