Skip to main content
QUICK REVIEW

[Paper Review] TRON: Transformer Neural Network Acceleration with Non-Coherent Silicon Photonics

Salma Afifi, Febin Sunny|arXiv (Cornell University)|Mar 22, 2023
Neural Networks and Reservoir Computing4 citations
TL;DR

TRON proposes the first silicon photonic neural network accelerator for Transformer models like BERT and Vision Transformers, leveraging non-coherent silicon photonics to achieve at least 14x higher throughput and 8x better energy efficiency than state-of-the-art electronic accelerators, enabling high-performance inference for NLP and computer vision workloads.

ABSTRACT

Transformer neural networks are rapidly being integrated into state-of-the-art solutions for natural language processing (NLP) and computer vision. However, the complex structure of these models creates challenges for accelerating their execution on conventional electronic platforms. We propose the first silicon photonic hardware neural network accelerator called TRON for transformer-based models such as BERT, and Vision Transformers. Our analysis demonstrates that TRON exhibits at least 14x better throughput and 8x better energy efficiency, in comparison to state-of-the-art transformer accelerators.

Motivation & Objective

  • Address the growing computational demands of Transformer-based models in NLP and computer vision.
  • Overcome the limitations of electronic accelerators in handling the high-precision, high-bandwidth computations of attention mechanisms.
  • Design a photonic hardware accelerator tailored for the unique computational patterns of Transformers.
  • Achieve significant improvements in throughput and energy efficiency compared to state-of-the-art electronic accelerators.
  • Demonstrate feasibility and performance gains of non-coherent silicon photonics for large-scale AI inference.

Proposed method

  • Design a photonic accelerator architecture optimized for the matrix multiplication and softmax operations central to self-attention mechanisms.
  • Utilize non-coherent silicon photonics to reduce system complexity and cost compared to coherent photonic systems.
  • Implement a hybrid electronic-photonic processing pipeline to handle control logic and data flow coordination.
  • Leverage wavelength-division multiplexing to parallelize multiple attention heads and compute paths.
  • Integrate photonic interconnects for high-bandwidth, low-latency communication between processing elements.
  • Optimize the mapping of attention computation to photonic domain operations to minimize conversion losses and latency.

Experimental results

Research questions

  • RQ1Can non-coherent silicon photonics enable efficient acceleration of Transformer neural networks?
  • RQ2How does a photonic accelerator compare to electronic counterparts in terms of throughput and energy efficiency for attention-based models?
  • RQ3What architectural and signal processing innovations are required to make photonic inference viable for real-world Transformers?
  • RQ4To what extent can photonic hardware reduce the latency and energy consumption of self-attention computation?
  • RQ5What are the performance trade-offs of using non-coherent optics in a scalable AI accelerator?

Key findings

  • TRON achieves at least 14x higher throughput compared to state-of-the-art electronic transformer accelerators.
  • The proposed accelerator demonstrates 8x better energy efficiency than existing electronic solutions.
  • Non-coherent silicon photonics enables a scalable and cost-effective photonic accelerator design without requiring complex phase stabilization.
  • The architecture effectively handles the high-precision, high-bandwidth requirements of attention mechanisms in BERT and Vision Transformers.
  • The performance gains are attributed to the parallelism and low-latency interconnects inherent in photonic domain computation.
  • The results validate the feasibility of using silicon photonics for accelerating modern deep learning workloads beyond traditional electronic limits.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.