[Paper Review] From pixels to planning: scale-free active inference
This paper introduces renormalising generative models (RGMs), a scale-invariant, discrete state-space framework that extends partially observed Markov decision processes by treating paths as latent variables. Using variational inference and renormalisation group principles, RGMs enable end-to-end learning of compositional, hierarchical representations across space and time—demonstrated through image classification, video/music generation, and Atari-like environment planning—achieving scale-free, hierarchical inference without architectural bias.
This paper describes a discrete state-space model -- and accompanying methods -- for generative modelling. This model generalises partially observed Markov decision processes to include paths as latent variables, rendering it suitable for active inference and learning in a dynamic setting. Specifically, we consider deep or hierarchical forms using the renormalisation group. The ensuing renormalising generative models (RGM) can be regarded as discrete homologues of deep convolutional neural networks or continuous state-space models in generalised coordinates of motion. By construction, these scale-invariant models can be used to learn compositionality over space and time, furnishing models of paths or orbits; i.e., events of increasing temporal depth and itinerancy. This technical note illustrates the automatic discovery, learning and deployment of RGMs using a series of applications. We start with image classification and then consider the compression and generation of movies and music. Finally, we apply the same variational principles to the learning of Atari-like games.
Motivation & Objective
- To develop a scale-invariant, hierarchical generative model that supports long-range temporal and spatial compositionality.
- To unify perception and planning under active inference by treating trajectories as latent variables in a discrete state-space framework.
- To enable automatic discovery and learning of compositional structures in sensory data (images, video, music, games) without architectural inductive bias.
- To demonstrate the applicability of the framework across diverse domains—from static image classification to dynamic, sequential decision-making in Atari-like environments.
- To establish a formal link between deep learning architectures and active inference through renormalisation group-inspired hierarchical abstraction.
Proposed method
- Proposes a discrete state-space model where latent variables represent paths or trajectories, extending partially observed Markov decision processes.
- Employs variational inference with free energy minimisation to perform approximate Bayesian inference over latent paths and states.
- Applies renormalisation group (RG) principles to hierarchically coarse-grain representations across spatial and temporal scales.
- Uses a hierarchical, multi-level generative model that learns compositional features through iterative coarse-graining and reconstruction.
- Integrates the framework into a unified active inference engine that supports both perception and planning via path inference.
- Leverages generalised coordinates of motion to model continuous dynamics in a discrete, scale-invariant manner.
Experimental results
Research questions
- RQ1Can a single, scale-free generative model learn hierarchical representations across diverse modalities such as images, video, music, and game environments?
- RQ2How can path or trajectory inference be embedded as a latent variable in a generative model to support both perception and planning?
- RQ3To what extent can renormalisation group principles be used to construct hierarchical, compositionally structured representations without architectural bias?
- RQ4Can the same variational inference framework support both generative modeling and decision-making in dynamic environments?
- RQ5How does the model’s scale-invariance enable robust learning and generalisation across different temporal and spatial scales?
Key findings
- The RGM framework successfully learns hierarchical, compositional representations from raw pixels, enabling end-to-end image classification with minimal inductive bias.
- The model achieves effective compression and generation of video and music sequences by learning long-range temporal dependencies through path inference.
- In Atari-like environments, the model learns to plan by inferring optimal action sequences (paths) that minimise free energy, demonstrating active inference in decision-making.
- The hierarchical structure enables scale-free reasoning, where representations at coarser levels capture abstract, long-horizon dynamics while finer levels retain local detail.
- The framework demonstrates robustness to distributional shifts and generalisation across tasks due to its intrinsic invariance to scale and resolution.
- The use of renormalisation group principles allows for systematic abstraction across multiple levels of temporal and spatial resolution, mimicking deep network hierarchies without fixed architecture.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.