[Paper Review] NeuroBench: A Framework for Benchmarking Neuromorphic Computing Algorithms and Systems
NeuroBench introduces a dual-track benchmark framework for neuromorphic computing, with an algorithm track for hardware-independent evaluation and a system track for hardware deployment, plus baseline results for several tasks.
Neuromorphic computing shows promise for advancing computing efficiency and capabilities of AI applications using brain-inspired principles. However, the neuromorphic research field currently lacks standardized benchmarks, making it difficult to accurately measure technological advancements, compare performance with conventional methods, and identify promising future research directions. Prior neuromorphic computing benchmark efforts have not seen widespread adoption due to a lack of inclusive, actionable, and iterative benchmark design and guidelines. To address these shortcomings, we present NeuroBench: a benchmark framework for neuromorphic computing algorithms and systems. NeuroBench is a collaboratively-designed effort from an open community of researchers across industry and academia, aiming to provide a representative structure for standardizing the evaluation of neuromorphic approaches. The NeuroBench framework introduces a common set of tools and systematic methodology for inclusive benchmark measurement, delivering an objective reference framework for quantifying neuromorphic approaches in both hardware-independent (algorithm track) and hardware-dependent (system track) settings. In this article, we outline tasks and guidelines for benchmarks across multiple application domains, and present initial performance baselines across neuromorphic and conventional approaches for both benchmark tracks. NeuroBench is intended to continually expand its benchmarks and features to foster and track the progress made by the research community.
Motivation & Objective
- Provide a formal, inclusive benchmarking framework for neuromorphic computing to enable fair comparison across diverse approaches.
- Introduce two cohesive tracks (algorithm and system) to cover both software/design and deployed hardware perspectives.
- Offer a community-driven, iterative benchmark suite with open tooling and leaderboards to track progress.
- Present baseline results on key neuromorphic benchmarks to guide future research in algorithmic efficiency and hardware deployment.
Proposed method
- Define hardware-independent (algorithm) and hardware-dependent (system) benchmark tracks.
- Specify a common harness and metrics to enable fair, repeatable evaluation across diverse neuromorphic approaches.
- Propose four algorithm-track benchmarks (few-shot continual learning, event-based object detection, motor cortical decoding, chaotic forecasting) with complexity and correctness metrics.
- Provide baseline algorithm results comparing ANN and SNN approaches on the FSCIL keyword task.
- Outline system-track protocols to evaluate real-world speed and efficiency of neuromorphic hardware on applicable workloads.

Experimental results
Research questions
- RQ1How can neuromorphic research be standardized to enable fair, inclusive benchmarking across diverse approaches?
- RQ2What are the key metrics that capture algorithmic complexity, correctness, and hardware-agnostic performance for neuromorphic methods?
- RQ3How do hardware-independent algorithm benchmarks inform subsequent system-track hardware design and deployment strategies?
- RQ4What baseline performances do current ANN and SNN approaches achieve on representative neuromorphic tasks under NeuroBench?
- RQ5How can NeuroBench evolve to accommodate new modalities and tasks over time?
Key findings
- NeuroBench defines a two-track framework (algorithm and system) with a common harness that enables hardware-agnostic comparisons and cross-stack influence between tracks.
- Algorithm-track benchmarks capture both correctness and complexity metrics, including footprint, model execution rate, connection sparsity, activation sparsity, and synaptic operations.
- In the v1.0 algorithm track, four benchmarks were established: FSCIL, event-camera object detection, non-human primate motor prediction, and chaotic function prediction.
- Baseline comparisons show M5 ANN and SNN baselines for the FSCIL task with differing footprints, execution rates, and synaptic operation profiles, illustrating trade-offs between accuracy and efficiency.
- The FSCIL baselines achieved base accuracies of 97.09% (ANN) and 93.48% (SNN) on base classes, with notable differences in footprint, execution rate, sparsity, and synaptic operations.
- The framework is open-source and designed for ongoing community-driven expansion, including potential additions of data modalities (e.g., IMU) and closed-loop tasks.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.