[Paper Review] SEE-MCAM: Scalable Multi-bit FeFET Content Addressable Memories for Energy Efficient Associative Search
This paper proposes SEE-MCAM, a scalable multi-bit FeFET-based content addressable memory (CAM) that leverages the multi-level cell (MLC) capability of ferroelectric FETs (FeFETs) to enable high-density, energy-efficient associative search. By designing NOR- and NAND-type 2FeFET-1T and 2FeFET-2T CAM architectures, SEE-MCAM achieves 3 bits per cell, 8% area per bit compared to CMOS CAM, 9.8× better energy efficiency, and 1.6× lower search latency than CMOS CAM, with up to 3× speedup and energy efficiency gain over GPU implementations in quantized hyperdimensional computing workloads.
In this work, we propose SEE-MCAM, scalable and compact multi-bit CAM (MCAM) designs that utilize the three-terminal ferroelectric FET (FeFET) as the proxy. By exploiting the multi-level-cell characteristics of FeFETs, our proposed SEE-MCAM designs enable multi-bit associative search functions and achieve better energy efficiency and performance than existing FeFET-based CAM designs. We validated the functionality of our proposed designs by achieving 3 bits per cell CAM functionality, resulting in 3x improvement in storage density. The area per bit of the proposed SEE-MCAM cell is 8% of the conventional CMOS CAM. We thoroughly investigated the scalability and robustness of the proposed design. Evaluation results suggest that the proposed 2FeFET-1T SEE-MCAM achieves 9.8x more energy efficiency and 1.6x less search latency compared to the CMOS CAM, respectively. When compared to existing MCAM designs, the proposed SEE-MCAM can achieve 8.7x and 4.9x more energy efficiency than ReRAM-based and FeFET-based MCAMs, respectively. Benchmarking results show that our approach provides up to 3 orders of magnitude improvement in speedup and energy efficiency over a GPU implementation in accelerating a novel quantized hyperdimensional computing (HDC) application.
Motivation & Objective
- To address the memory wall bottleneck in AI workloads by enabling high-density, low-energy associative search within memory.
- To overcome the limitations of existing binary and multi-bit CAMs based on NVMs, which are constrained by single-level cell (SLC) operation and suffer from high area, energy, or sensitivity to device variation.
- To exploit the multi-level cell (MLC) characteristics of FeFETs to enable scalable, compact, and robust multi-bit CAM (MCAM) designs with improved energy efficiency and performance.
- To validate the feasibility and superiority of the proposed FeFET-based MCAM in real-world AI workloads, particularly quantized hyperdimensional computing (HDC).
Proposed method
- Designing a 2FeFET multi-level cell (MLC) structure using the MIBO (multi-level inverter-based output) concept to store multiple bits per FeFET.
- Proposing two CAM architectures: NOR-type 2FeFET-1T and NAND-type 2FeFET-2T, to support multi-bit associative search with minimal area and energy overhead.
- Implementing a non-linear quantization scheme to map hypervector elements to 3-bit values, enabling compatibility with the 3-bit-per-cell SEE-MCAM design.
- Integrating the SEE-MCAM into a quantized hyperdimensional computing (HDC) framework to benchmark performance and energy efficiency against GPU and other CAM-based implementations.
- Using a 16T CMOS CAM as baseline for comparison, and evaluating performance via simulation and hardware-aware benchmarks on real datasets (ISOLET, UCIHAR, PAMAP).
- Employing Nvidia System Management Interface and PyTorch profiler for accurate power and delay measurement in GPU and SEE-MCAM-based HDC workloads.
Experimental results
Research questions
- RQ1Can FeFET-based multi-bit CAMs achieve higher storage density and energy efficiency than existing binary and multi-bit CAMs?
- RQ2How can the multi-level cell (MLC) property of FeFETs be effectively exploited to enable scalable, low-energy multi-bit associative search?
- RQ3What are the performance and energy efficiency gains of SEE-MCAM over CMOS-based CAMs and other NVM-based CAMs in real AI workloads?
- RQ4To what extent does SEE-MCAM improve accuracy and efficiency in quantized hyperdimensional computing (HDC) applications compared to binary CAMs and GPU implementations?
Key findings
- The proposed 2FeFET-1T SEE-MCAM achieves 9.8× higher energy efficiency and 1.6× lower search latency compared to 16T CMOS CAM.
- SEE-MCAM achieves 3 bits per cell functionality, resulting in a 3× improvement in storage density over binary CAMs.
- The area per bit of the SEE-MCAM cell is only 8% of that of conventional CMOS CAM, demonstrating significant area scaling.
- Compared to ReRAM-based and FeFET-based MCAMs, SEE-MCAM achieves 8.7× and 4.9× higher energy efficiency, respectively.
- In quantized HDC workloads, SEE-MCAM-based implementations deliver up to 3 orders of magnitude improvement in both speedup and energy efficiency over GPU (GTX 1080ti) implementations.
- The 3-bit SEE-MCAM achieves 2.41% higher accuracy than the 2-bit version due to increased hypervector dimensionality, and 2.26% higher accuracy than a binary COSIME-based implementation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.