Skip to main content
QUICK REVIEW

[论文解读] 3D-aCortex: An Ultra-Compact Energy-Efficient Neurocomputing Platform Based on Commercial 3D-NAND Flash Memories

Mohammad Bavandpour, Shubham Sahay|arXiv (Cornell University)|Aug 7, 2019
Advanced Memory and Neural Computing被引用 5
一句话总结

本文提出3D-aCortex,一种超紧凑、能效极高的神经形态推理处理器,通过时间域编码的向量-矩阵乘法(VMM)技术,利用商用3D-NAND闪存块实现存内计算。该设计在不修改现有64层3D-NAND技术的前提下,实现了前所未有的能效比(约10 fJ/Op)和存储密度(4.34 MB/mm²),峰值性能达70.43 TOps/J和10.66 TOps/s。

ABSTRACT

The first contribution of this paper is the development of extremely dense, energy-efficient mixed-signal vector-by-matrix-multiplication (VMM) circuits based on the existing 3D-NAND flash memory blocks, without any need for their modification. Such compatibility is achieved using time-domain-encoded VMM design. Our detailed simulations have shown that, for example, the 5-bit VMM of 200-element vectors, using the commercially available 64-layer gate-all-around macaroni-type 3D-NAND memory blocks designed in the 55-nm technology node, may provide an unprecedented area efficiency of 0.14 um2/byte and energy efficiency of ~10 fJ/Op, including the input/output and other peripheral circuitry overheads. Our second major contribution is the development of 3D-aCortex, a multi-purpose neuromorphic inference processor that utilizes the proposed 3D-VMM blocks as its core processing units. We have performed rigorous performance simulations of such a processor on both circuit and system levels, taking into account non-idealities such as drain-induced barrier lowering, capacitive coupling, charge injection, parasitics, process variations, and noise. Our modeling of the 3D-aCortex performing several state-of-the-art neuromorphic-network benchmarks has shown that it may provide the record-breaking storage efficiency of 4.34 MB/mm2, the peak energy efficiency of 70.43 TOps/J, and the computational throughput up to 10.66 TOps/s. The storage efficiency can be further improved seven-fold by aggressively sharing VMM peripheral circuits at the cost of slight decrease in energy efficiency and throughput.

研究动机与目标

  • 设计一种基于现有商用3D-NAND闪存、无需硬件修改的高能效且紧凑的神经计算平台。
  • 通过在3D-NAND闪存块中采用时间域编码,实现向量-矩阵乘法(VMM)的极致面积与能效优化。
  • 开发一种多功能神经形态推理处理器3D-aCortex,具备高吞吐量与高存储密度,同时考虑非理想器件效应的影响。
  • 在真实非理想条件(如工艺失配、噪声和寄生效应)下评估系统级性能。
  • 在紧凑、可扩展的神经形态架构中实现创纪录的能效比与存储密度。

提出的方法

  • 采用时间域编码的VMM技术,在不改变结构的前提下,实现现有3D-NAND闪存块中的混合信号计算。
  • VMM电路基于55-nm工艺节点制造的64层全栅环绕式圆筒型3D-NAND闪存块实现。
  • 在电路与系统两级层面建模并仿真非理想器件效应,如漏极诱导势垒降低、电容耦合、电荷注入及工艺失配。
  • 设计了一种多核神经形态处理器架构3D-aCortex,以3D-VMM模块作为核心处理单元,支持推理工作负载。
  • 探索VMM外围电路共享机制,以进一步提升存储效率,代价为轻微的吞吐量与能效比折损。
  • 在多个基准测试上开展全面仿真,评估性能、能效比与面积效率。

实验结果

研究问题

  • RQ1是否能在不修改结构的前提下,高效地在商用3D-NAND闪存中实现存内向量-矩阵乘法?
  • RQ2在3D-NAND闪存中采用时间域编码VMM,可实现的能效与面积效率达到何种水平?
  • RQ33D-aCortex在真实非理想条件(如工艺失配、噪声与寄生效应)下的表现如何?
  • RQ4基于3D-NAND的神经形态处理器可实现的存储密度与计算吞吐量是多少?
  • RQ5外围电路共享能否显著提升存储效率而不显著降低核心性能指标?

主要发现

  • 所提出的3D-VMM实现0.14 µm²/byte的面积效率与约10 fJ/Op的能效比,包含I/O与外围电路开销。
  • 3D-aCortex利用64层3D-NAND闪存块,实现创纪录的4.34 MB/mm²存储密度。
  • 该处理器展现出70.43 TOps/J的峰值能效比与最高达10.66 TOps/s的计算吞吐量。
  • VMM外围电路的激进共享机制使存储效率最高提升七倍,尽管伴随轻微的能效比与吞吐量下降。
  • 仿真结果表明,在真实非理想条件下,该设计在多个先进神经形态网络基准测试中均表现出稳健性能。
  • 该设计与现有商用3D-NAND闪存技术完全兼容,具备即插即用的可扩展性与集成能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。