[Paper Review] A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms
This survey reviews and classifies the latest deep learning accelerators for heterogeneous HPC platforms, covering GPUs, TPUs, FPGA/ASIC NPUs, open-hardware RISC-V co-processors, and emerging memory and computing paradigms. It provides a consolidated view of architectures, technologies, and future challenges.
Recent trends in deep learning (DL) have made hardware accelerators essential for various high-performance computing (HPC) applications, including image classification, computer vision, and speech recognition. This survey summarizes and classifies the most recent developments in DL accelerators, focusing on their role in meeting the performance demands of HPC applications. We explore cutting-edge approaches to DL acceleration, covering not only GPU- and TPU-based platforms but also specialized hardware such as FPGA- and ASIC-based accelerators, Neural Processing Units, open hardware RISC-V-based accelerators, and co-processors. This survey also describes accelerators leveraging emerging memory technologies and computing paradigms, including 3D-stacked Processor-In-Memory, non-volatile memories like Resistive RAM and Phase Change Memories used for in-memory computing, as well as Neuromorphic Processing Units, and Multi-Chip Module-based accelerators. Furthermore, we provide insights into emerging quantum-based accelerators and photonics. Finally, this survey categorizes the most influential architectures and technologies from recent years, offering readers a comprehensive perspective on the rapidly evolving field of deep learning acceleration.
Motivation & Objective
- Provide a comprehensive overview of influential DL accelerator architectures for HPC.
- Classify accelerators by hardware type, memory tech, and computing paradigm to highlight similarities and differences.
- Explain dataflow and memory-reuse strategies that impact energy and performance in DL accelerators.
- Summarize emerging technologies and future challenges in DL accelerator research.
- Offer a reference point for researchers and practitioners designing HPC systems with DL workloads.
Proposed method
- Classifies accelerators using representative features to compare architectures.
- Reviews and references about ~230 works on DL acceleration.
- Covers GPU, TPU, FPGA, ASIC NPUs, RISC-V open-hardware accelerators, and co-processors.
- Describes emerging paradigms such as 3D-stacked PIM, non-volatile memories (RRAM/PCM), Neuromorphic units, and Multi-Chip Modules.
- Discusses future trends including quantum accelerators and photonics.
Experimental results
Research questions
- RQ1What are the most influential architectures and technologies enabling DL acceleration for HPC workloads?
- RQ2How do GPU/TPU, FPGA/ASIC NPUs, and open-hardware accelerators compare in terms of performance, energy efficiency, and flexibility for DL tasks?
- RQ3What role do emerging memory technologies and computing paradigms (e.g., PIM, RRAM/PCM, neuromorphic) play in DL acceleration?
- RQ4What are the key challenges and future directions for DL accelerators on heterogeneous HPC platforms?
- RQ5How can DL accelerators be organized and classified to provide a coherent perspective for researchers and practitioners?
Key findings
- DL accelerators span GPUs, TPUs, FPGAs, ASIC NPUs, open-hardware RISC-V co-processors, and memory/processing paradigms.
- Emerging technologies include 3D-stacked PIM, RRAM, PCM, Neuromorphic Processing Units, and Multi-Chip Modules.
- The survey analyzes and references about 230 works on DL acceleration to offer a comprehensive perspective.
- It discusses future challenges such as quantum accelerators and photonics for DL workloads.
- A structured classification helps compare architectures and identify trade-offs across speed, energy, and flexibility.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.