[Paper Review] Learning identifiable and interpretable latent models of high-dimensional neural activity using pi-VAE
pi-VAE integrates identifiable VAE concepts with task-variable conditioning to learn interpretable, identifiable latent structures in high-dimensional neural data, improving fit and latent interpretability over standard VAEs and tuning-curve models.
The ability to record activities from hundreds of neurons simultaneously in the brain has placed an increasing demand for developing appropriate statistical techniques to analyze such data. Recently, deep generative models have been proposed to fit neural population responses. While these methods are flexible and expressive, the downside is that they can be difficult to interpret and identify. To address this problem, we propose a method that integrates key ingredients from latent models and traditional neural encoding models. Our method, pi-VAE, is inspired by recent progress on identifiable variational auto-encoder, which we adapt to make appropriate for neuroscience applications. Specifically, we propose to construct latent variable models of neural activity while simultaneously modeling the relation between the latent and task variables (non-neural variables, e.g. sensory, motor, and other externally observable states). The incorporation of task variables results in models that are not only more constrained, but also show qualitative improvements in interpretability and identifiability. We validate pi-VAE using synthetic data, and apply it to analyze neurophysiological datasets from rat hippocampus and macaque motor cortex. We demonstrate that pi-VAE not only fits the data better, but also provides unexpected novel insights into the structure of the neural codes.
Motivation & Objective
- Motivate the need for models that are both flexible and interpretable for high-dimensional neural data.
- Develop a generative framework that jointly models latent structure and its relation to task variables.
- Achieve identifiability and interpretability by incorporating a label prior and Poisson observation noise.
- Demonstrate that pi-VAE yields better data fit and more interpretable latent structure than baselines on neural datasets.
Proposed method
- Define a generative model p_theta(x,z|u) = p_f(x|z) p_{T,lambda}(z|u).
- Model the label prior p_{T,lambda}(z|u) as conditionally independent exponential-family distributions with neural-network parameterized natural parameters lambda(u).
- Represent p_f(x|z) as Poisson observations with firing rate f(z) implemented via a generalized invertible flow (GIN) to handle high-dimensional outputs.
- Extend GIN to map from m-dimensional z to n-dimensional x, ensuring f is injective.
- Adopt an identifiable VAE-like inference scheme with q(z|x,u) ∝ q_phi(z|x) p_{T,lambda}(z|u), using Gaussian approximations for tractable training.
- Prove identifiability under mild conditions and validate with synthetic data and electrophysiology datasets.
Experimental results
Research questions
- RQ1Can pi-VAE recover interpretable and identifiable latent structure in high-dimensional neural recordings while leveraging task variables as labels?
- RQ2Does incorporating a label prior improve model fit and latent disentanglement compared to standard VAEs or tuning-curve approaches?
- RQ3What kind of latent geometry emerges in real neural data (e.g., hippocampus CA1 and M1/PMd) when using pi-VAE?
- RQ4How does pi-VAE perform in decoding and encoding tasks relative to traditional methods on neural datasets?
Key findings
- pi-VAE yields better fit to held-out neural data than VAE and tuning-curve models, as measured by higher marginal log-likelihood on test data.
- pi-VAE decodes task-relevant variables (e.g., reaching direction) more accurately than tuning-curve models, especially early in movement.
- Latent spaces from pi-VAE show interpretable geometry, separating direction information and temporal dynamics (e.g., in monkey reaching data, first two latent dims encode direction, others capture trajectory evolution).
- On rat CA1 data, pi-VAE reveals latent manifolds that align with track geometry and show directional separation consistent with place-cell properties, with temporal structure linked to theta rhythms (~10 Hz).
- Latent structures from pi-VAE are more disentangled and physically meaningful than those from alternative methods (UMAP, PfLDS, LFADS, demixed PCA).
- The label-prior and identifiable framework enable posterior inferences p(z|x,u) that are interpretable up to identifiable transformations, supporting scientific insights into neural codes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.