Skip to main content
QUICK REVIEW

[Paper Review] AI and ML Accelerator Survey and Trends

Albert Reuther, Peter Michaleas|arXiv (Cornell University)|Oct 8, 2022
Advanced Memory and Neural Computing4 citations
TL;DR

This paper surveys publicly announced AI and ML accelerators from the past three years, focusing on peak performance and power efficiency across diverse architectures including DNN/CNN inference processors, neuromorphic chips (e.g., Loihi), photonic processors, and memristor-based systems. It presents updated scatter plots of performance vs. power, identifies emerging trends in accelerator design, and analyzes trade-offs in computational effectiveness for embedded and data center applications.

ABSTRACT

This paper updates the survey of AI accelerators and processors from past three years. This paper collects and summarizes the current commercial accelerators that have been publicly announced with peak performance and power consumption numbers. The performance and power values are plotted on a scatter graph, and a number of dimensions and observations from the trends on this plot are again discussed and analyzed. Two new trends plots based on accelerator release dates are included in this year's paper, along with the additional trends of some neuromorphic, photonic, and memristor-based inference accelerators.

Motivation & Objective

  • To provide a comprehensive, up-to-date survey of publicly announced AI and ML accelerators with measured peak performance and power consumption.
  • To analyze trends in accelerator design, particularly focusing on inference workloads for low-power embedded systems and high-performance data centers.
  • To evaluate emerging non-traditional accelerator technologies such as neuromorphic, photonic, and memristor-based processors.
  • To compare computational effectiveness across different precision formats (int8, bfloat16, fp16) and architectural approaches.
  • To support defense and national security applications by identifying accelerators suitable for size, weight, and power-constrained environments.

Proposed method

  • Collected and compiled publicly available performance and power data from 150+ AI accelerators released between 2020 and 2023.
  • Plotted peak performance (TOPS) against power consumption (W) to visualize performance-per-watt trade-offs across accelerator types.
  • Grouped accelerators by application domain (edge, data center, embedded) and technology type (GPU, TPU, FPGA, neuromorphic, photonic, memristor).
  • Analyzed trends using time-ordered scatter plots based on accelerator release dates to track evolution in performance and efficiency.
  • Evaluated support for mixed-precision inference (int8, bfloat16, fp16) and assessed their impact on computational throughput.
  • Included case studies and references to emerging technologies like Loihi (neuromorphic), photonic co-processors, and Knowm memristors.

Experimental results

Research questions

  • RQ1What are the current performance and power efficiency trends in AI accelerators for deep neural network inference?
  • RQ2How do neuromorphic, photonic, and memristor-based accelerators compare in performance and energy efficiency to conventional GPU and TPU architectures?
  • RQ3What are the dominant design trade-offs in modern AI accelerators across different application domains (edge vs. data center)?
  • RQ4How has the development cycle for AI accelerators evolved over the past three years, and what factors influence time-to-market?
  • RQ5What role do mixed-precision formats (int8, bfloat16) play in determining the computational effectiveness of AI accelerators?

Key findings

  • Neuromorphic accelerators like Loihi 2 demonstrate ultra-low power consumption (under 100 mW) for spiking neural network inference, enabling real-time processing in energy-constrained environments.
  • Photonic accelerators, such as those using silicon photonics and Fourier optics, achieve theoretical performance gains of up to 100x over GPUs for specific operations like convolution and matrix multiplication.
  • Memristor-based accelerators show promise for in-memory computing, with Knowm devices achieving sub-microsecond switching speeds and low energy per operation in proof-of-concept studies.
  • The performance-per-watt of modern AI accelerators has improved by 2.5x–3x over the past three years, with many new chips achieving >100 TOPS/W in int8 inference.
  • Accelerators based on dataflow architectures (e.g., MIT’s Eyeriss) show superior energy efficiency for CNN inference, particularly in edge applications.
  • Despite progress, development cycles for new accelerators remain long—typically 2–4 years—indicating sustained R&D investment is required for innovation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.