Skip to main content
QUICK REVIEW

[Paper Review] Neural Latents Benchmark '21: Evaluating latent variable models of neural population activity

Felix Pei, Joel Ye|arXiv (Cornell University)|Sep 9, 2021
Neural dynamics and brain functionNeuroscience88 references44 citations
TL;DR

Introduces Neural Latents Benchmark (NLB) ’21 to standardize evaluation of unsupervised latent variable models (LVMs) on neural population data across diverse brain regions, tasks, and dataset sizes, using co-smoothing as the primary metric and EvalAI-hosted benchmarks.

ABSTRACT

Advances in neural recording present increasing opportunities to study neural activity in unprecedented detail. Latent variable models (LVMs) are promising tools for analyzing this rich activity across diverse neural systems and behaviors, as LVMs do not depend on known relationships between the activity and external experimental variables. However, progress with LVMs for neuronal population activity is currently impeded by a lack of standardization, resulting in methods being developed and compared in an ad hoc manner. To coordinate these modeling efforts, we introduce a benchmark suite for latent variable modeling of neural population activity. We curate four datasets of neural spiking activity from cognitive, sensory, and motor areas to promote models that apply to the wide variety of activity seen across these areas. We identify unsupervised evaluation as a common framework for evaluating models across datasets, and apply several baselines that demonstrate benchmark diversity. We release this benchmark through EvalAI. http://neurallatents.github.io

Motivation & Objective

  • Motivate standardized evaluation of latent variable models (LVMs) for neural population activity.
  • Provide curated datasets spanning motor, sensory, and cognitive regions and varying dataset sizes.
  • Define unsupervised evaluation framework with a robust primary metric (co-smoothing) and complementary metrics.
  • Offer a reproducible pipeline with Neurodata Without Borders-formatted data and EvalAI-based evaluation.
  • Baseline comparisons to establish performance benchmarks across model types.

Proposed method

  • Curate four diverse neural spiking datasets (MC_Maze, MC_RTT, Area2_Bump, DMFC_RSG) with varying task demands.
  • Adopt an unsupervised evaluation framework (co-smoothing) to predict held-out neural activity.
  • Provide secondary metrics: PSTH match, forward prediction, and behavioral decoding where applicable.
  • Compare five baseline LVM approaches (Smoothed spikes, GPFA, SLDS, AutoLFADS, Neural Data Transformer) to establish performance baselines.
  • Utilize train/val/test splits with private test data on EvalAI to prevent overfitting and hyperparameter hacking.

Experimental results

Research questions

  • RQ1How well can unsupervised latent variable models describe neural population activity across diverse brain regions and behaviors?
  • RQ2What is the relative performance of different LVM families (linear, nonlinear, deep) on held-out neuron/time predictions (co-smoothing) and secondary metrics?
  • RQ3How does model performance scale with dataset size and across datasets with varying numbers of neurons and firing rates?
  • RQ4Can the benchmark identify regimes where deep models (AutoLFADS, NDT) outperform traditional baselines, and where simpler models suffice?
  • RQ5How do evaluation metrics (co-smoothing vs. PSTH matching vs. forward prediction vs. behavioral decoding) align or diverge across datasets?

Key findings

  • Co-smoothing is generally achievable across datasets, with deep models often outperforming baselines.
  • Deep networks (AutoLFADS, NDT) show strong performance across multiple datasets, especially in more cognitive areas (DMFC_RSG) and larger datasets.
  • Performance advantages of deep models are dataset-dependent, with some datasets showing smaller gaps.
  • Benchmarked datasets scale from MC_Maze-L/M/S to MC_Maze and other tasks, illustrating dataset-size effects on LVM evaluation.
  • GPFA and SLDS show variable performance across metrics, highlighting the importance of choosing appropriate evaluation criteria.
  • Across datasets, deep models tend to offer the most consistent gains in co-smoothing performance.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.