[论文解读] A Time-domain Analog Weighted-sum Calculation Model for Extremely Low Power VLSI Implementation of Multi-layer Neural Networks
本文提出一种基于积分-发放脉冲神经元的时间域模拟加权求和计算模型,用于实现超低功耗的多层神经网络VLSI芯片。通过利用亚阈值MOSFET实现高电阻的瞬态RC充电,该方法在250-nm CMOS工艺下实现了290 TOPS/W的能量效率——比当前最先进的数字AI处理器高出十倍以上,且在先进工艺下有望实现超过1,000 TOPS/W的性能。
A time-domain analog weighted-sum calculation model is proposed based on an integrate-and-fire-type spiking neuron model. The proposed calculation model is applied to multi-layer feedforward networks, in which weighted summations with positive and negative weights are separately performed in each layer and summation results are then fed into the next layers without their subtraction operation. We also propose very large-scale integrated (VLSI) circuits to implement the proposed model. Unlike the conventional analog voltage or current mode circuits, the time-domain analog circuits use transient operation in charging/discharging processes to capacitors. Since the circuits can be designed without operational amplifiers, they can operate with extremely low power consumption. However, they have to use very high resistance devices on the order of G$ m Ω$. We designed a proof-of-concept (PoC) CMOS VLSI chip to verify weighted-sum operation with the same weights and evaluated it by post-layout circuit simulation using 250-nm fabrication technology. High resistance operation was realized by using the subthreshold operation region of MOS transistors. Simulation results showed that energy efficiency for the weighted-sum calculation was 290~TOPS/W, more than one order of magnitude higher than that in state-of-the-art digital AI processors, even though the minimum width of interconnection used in the PoC chip was several times larger than that in such digital processors. If state-of-the-art VLSI technology is used to implement the proposed model, an energy efficiency of more than 1,000~TOPS/W will be possible. For practical applications, development of emerging analog memory devices such as ferroelectric-gate FETs is necessary.
研究动机与目标
- 实现面向边缘AI应用的多层神经网络在超低功耗VLSI中的实现。
- 克服深度神经网络中传统数字乘加(MAC)操作带来的高能耗问题。
- 开发一种无需运算放大器的时间域模拟计算方法,实现超低功耗运行。
- 通过250-nm工艺的原型CMOS VLSI芯片验证该方法的可行性。
- 识别该方法在扩展过程中面临的关键挑战,尤其是新兴模拟存储器件与低功耗比较器的需求。
提出的方法
- 该模型采用积分-发放脉冲神经元框架,其中输入脉冲的到达时间调制突触后电位的上升时间。
- 加权求和通过RC电路实现,时间常数由电阻与电容决定,其中电阻通过亚阈值MOSFET实现,以获得GΩ量级的阻值。
- 正负权重通过独立的RC路径并行计算,结果在不进行显式减法的情况下合并。
- 神经元部分使用比较器检测阈值穿越,ReLU激活通过简单的逻辑门实现,输出两个时间信号中较晚的一个。
- 设计并验证了一款基于250-nm工艺的原型CMOS VLSI芯片,采用版图后仿真进行验证。
- 该电路架构支持多层前馈网络,支持层间直接的脉冲时间传播。
实验结果
研究问题
- RQ1基于瞬态RC状态的时间域模拟计算,是否能显著降低神经网络中数字MAC操作的能量消耗?
- RQ2在模拟域中,如何高效地计算并合并正负权重,而无需执行减法操作?
- RQ3在CMOS VLSI中,使用亚阈值MOSFET实现高阻值元件,可达到何种能效水平?
- RQ4与传统数字AI处理器相比,该模型在能效和精度方面表现如何?
- RQ5在实际部署中,该方法面临的主要硬件瓶颈是什么——尤其是存储器与比较器设计方面?
主要发现
- 所提出的时域模拟加权求和模型在250-nm CMOS工艺的版图后仿真中实现了290 TOPS/W的能效。
- 该能效比当前最先进的数字AI处理器高出一个数量级以上。
- 突触电路每操作功耗约为3.34 fJ,神经元比较器功耗约为1.67 pJ,表明比较器是主要的功耗来源。
- 该方法避免使用运算放大器,仅依赖无源RC元件与开关,从而实现超低功耗运行。
- 若采用寄生电容更低的先进VLSI工艺,能效有望超过1,000 TOPS/W。
- 该方法对模拟非理想性具有鲁棒性,且在MNIST数据集上使用多层感知机测试时,其精度与数值计算相当。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。