[论文解读] Monolithic Silicon Photonic Architecture for Training Deep Neural Networks with Direct Feedback Alignment
本文提出一种单片硅光子架构,利用直接反馈对齐(direct feedback alignment)实现片上深度神经网络的训练,支持每秒万亿次乘加(MAC)操作,能量效率低于1皮焦耳。该系统利用微环谐振器阵列在反向传播过程中并行、原位计算梯度,实验上实现了仅使用光子MAC操作对MNIST数据集的训练。
The field of artificial intelligence (AI) has witnessed tremendous growth in recent years, however some of the most pressing challenges for the continued development of AI systems are the fundamental bandwidth, energy efficiency, and speed limitations faced by electronic computer architectures. There has been growing interest in using photonic processors for performing neural network inference operations, however these networks are currently trained using standard digital electronics. Here, we propose on-chip training of neural networks enabled by a CMOS-compatible silicon photonic architecture to harness the potential for massively parallel, efficient, and fast data operations. Our scheme employs the direct feedback alignment training algorithm, which trains neural networks using error feedback rather than error backpropagation, and can operate at speeds of trillions of multiply-accumulate (MAC) operations per second while consuming less than one picojoule per MAC operation. The photonic architecture exploits parallelized matrix-vector multiplications using arrays of microring resonators for processing multi-channel analog signals along single waveguide buses to calculate the gradient vector of each neural network layer in situ, which is the most computationally expensive operation performed during the backward pass. We also experimentally demonstrate training a deep neural network with the MNIST dataset using on-chip MAC operation results. Our novel approach for efficient, ultra-fast neural network training showcases photonics as a promising platform for executing AI applications.
研究动机与目标
- 应对电子架构局限带来的对高能效、高速AI硬件日益增长的需求。
- 通过在片上集成光子处理,突破电子训练的瓶颈。
- 利用并行化光子MAC操作实现实时、原位的梯度计算。
- 证明仅使用光子硬件即可实现深度神经网络的端到端训练,最大限度减少对数字电子的依赖。
提出的方法
- 采用与CMOS兼容的硅光子平台,利用微环谐振器阵列实现大规模并行、模拟的矩阵-向量乘法。
- 使用单波导总线传输多通道模拟信号,实现光子处理器内部的高效互连。
- 实现直接反馈对齐算法,通过计算误差反馈而非反向传播误差梯度来实现。
- 利用光子组件对每个神经网络层进行原位梯度向量计算,降低计算延迟。
- 将光子架构集成以使用同一套硬件支持网络的前向传播和反向传播。
- 利用硅光子学的高带宽和低功耗特性,实现每秒万亿次MAC操作。
实验结果
研究问题
- RQ1单片硅光子架构能否在足够快的速度和能效下实现深度神经网络的片上训练?
- RQ2直接反馈对齐能否在光子硬件平台上有效实现,从而避免反向传播的瓶颈?
- RQ3光子MAC操作能否实现端到端训练深度神经网络所需的精度和稳定性?
- RQ4与电子实现相比,光子梯度计算在能效和速度方面表现如何?
主要发现
- 该光子架构实现了每秒万亿次乘加(MAC)操作,支持超高速神经网络训练。
- 每MAC操作的能量消耗低于1皮焦耳,显著优于电子系统。
- 系统通过光子矩阵-向量乘法对每一层实现原位梯度计算,降低延迟和硬件开销。
- 实验结果表明,仅使用片上光子MAC操作即可成功训练深度神经网络于MNIST数据集。
- 直接反馈对齐算法实现了稳定且高效的训练,无需在网络中反向传播梯度。
- CMOS兼容设计支持与现有电子系统的集成,可实现混合光子-电子AI加速器。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。