[论文解读] Sub-1-us, Sub-20-nJ Pattern Classification in a Mixed-Signal Circuit Based on Embedded 180-nm Floating-Gate Memory Cell Arrays
本文提出一种采用嵌入式180-nm浮空栅存储器阵列的混合信号类脑电路,实现亚1-μs、亚20-nJ的模式分类。通过利用可调浮空栅单元实现精确的模拟向量-矩阵乘法,该原型在MNIST数据集上实现了94.65%的准确率,其速度和能效性能超过最先进的数字实现方案。
We have designed, fabricated, and successfully tested a prototype mixed-signal, 28x28-binary-input, 10-output, 3-layer neuromorphic network ("MLP perceptron"). It is based on embedded nonvolatile floating-gate cell arrays redesigned from a commercial 180-nm NOR flash memory. The arrays allow precise (~1%) individual tuning of all memory cells, having long-term analog-level retention and low noise. Each array performs a very fast and energy-efficient analog vector-by-matrix multiplication, which is the bottleneck for signal propagation in most neuromorphic networks. All functional components of the prototype circuit, including 2 synaptic arrays with 101,780 floating-gate synaptic cells, 74 analog neurons, and the peripheral circuitry for weight adjustment and I/O operations, have a total area below 1 mm^2. Its testing on the common MNIST benchmark set (at this stage, with a relatively low weight import precision) has shown a classification fidelity of 94.65%, close to the 96.2% obtained in simulation. The classification of one pattern takes less than 1 us time and ~20 nJ energy - both numbers much better than for digital implementations of the same task. Estimates show that this performance may be further improved using a better neuron design and a more advanced memory technology, leading to a >10^2 advantage in speed and a >10^4 advantage in energy efficiency over the state-of-the-art purely digital (GPU and custom) circuits, at classification of large, complex patterns.
研究动机与目标
- 开发一种用于实时模式分类的低功耗、高速类脑处理器,采用嵌入式非易失性存储器。
- 通过模拟存内计算技术,克服传统数字实现方案在神经网络推理中面临的能量和延迟瓶颈。
- 展示一种完全集成的混合信号电路,采用嵌入式浮空栅存储器阵列,实现可扩展、高能效的神经网络加速。
- 验证商用180-nm快闪技术在类脑系统中用于高精度、低噪声突触权重存储的可行性。
- 实现显著超越现有最先进数字加速器的性能指标,包括模式分类任务中的速度和能效。
提出的方法
- 采用嵌入式180-nm NOR快闪存储器阵列作为非易失性突触权重存储介质,经重新配置后实现约1%分辨率的精确模拟调节。
- 实现一个包含28×28二值输入、10个输出和74个模拟神经元的三层前馈MLP架构,所有组件均集成于单颗芯片上。
- 利用浮空栅单元在存储器内直接执行模拟向量-矩阵乘法,消除片外数据传输,降低延迟。
- 设计外围电路用于原位权重调节、输入/输出接口连接及信号调理,以支持实时运行。
- 在浮空栅单元中实现低噪声和长期模拟保持特性,确保突触权重值的稳定与可重复性。
- 将所有组件——包括101,780个突触单元和74个神经元——集成于总面积小于1 mm²的芯片内。
实验结果
研究问题
- RQ1嵌入式180-nm浮空栅存储器阵列能否被有效重用于类脑计算中的高精度、低噪声突触权重存储?
- RQ2与数字实现方案相比,浮空栅阵列中的模拟存内计算在多大程度上可降低模式分类任务中的能耗和延迟?
- RQ3采用该方法的完全集成混合信号类脑电路可实现的分类准确率及性能(延迟、能耗)如何?
- RQ4在当前原型基础上,通过神经元设计与存储技术的改进,性能将如何扩展?
- RQ5在单芯片上集成模拟神经元与存储阵列,能否实现亚微秒级分类速度与亚20-nJ能耗?
主要发现
- 原型实现每输入模式分类延迟小于1 μs,证明具备亚微秒级推理能力。
- 每次分类的能耗测量值约为20 nJ,显著低于现有最先进数字加速器。
- 在MNIST基准测试中,该电路实现94.65%的分类保真度,接近模拟中观察到的96.2%准确率。
- 系统在浮空栅单元中表现出稳定的长期保持特性与低噪声,支持约1%分辨率的精确模拟权重调节。
- 整个电路(包括101,780个突触单元和74个模拟神经元)集成于小于1 mm²的紧凑面积内。
- 性能估算表明,通过先进神经元设计与存储技术,有望实现速度提升100倍以上、能效提升超过10,000倍。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。