Skip to main content
QUICK REVIEW

[Paper Review] SEE-MCAM: Scalable Multi-bit FeFET Content Addressable Memories for Energy Efficient Associative Search

Shengxi Shou, Che-Kai Liu|arXiv (Cornell University)|Oct 7, 2023
Ferroelectric and Negative Capacitance Devices4 citations
TL;DR

This paper proposes SEE-MCAM, a scalable multi-bit FeFET-based content addressable memory (CAM) that leverages the multi-level cell (MLC) capability of ferroelectric FETs (FeFETs) to enable high-density, energy-efficient associative search. By designing NOR- and NAND-type 2FeFET-1T and 2FeFET-2T CAM architectures, SEE-MCAM achieves 3 bits per cell, 8% area per bit compared to CMOS CAM, 9.8× better energy efficiency, and 1.6× lower search latency than CMOS CAM, with up to 3× speedup and energy efficiency gain over GPU implementations in quantized hyperdimensional computing workloads.

ABSTRACT

In this work, we propose SEE-MCAM, scalable and compact multi-bit CAM (MCAM) designs that utilize the three-terminal ferroelectric FET (FeFET) as the proxy. By exploiting the multi-level-cell characteristics of FeFETs, our proposed SEE-MCAM designs enable multi-bit associative search functions and achieve better energy efficiency and performance than existing FeFET-based CAM designs. We validated the functionality of our proposed designs by achieving 3 bits per cell CAM functionality, resulting in 3x improvement in storage density. The area per bit of the proposed SEE-MCAM cell is 8% of the conventional CMOS CAM. We thoroughly investigated the scalability and robustness of the proposed design. Evaluation results suggest that the proposed 2FeFET-1T SEE-MCAM achieves 9.8x more energy efficiency and 1.6x less search latency compared to the CMOS CAM, respectively. When compared to existing MCAM designs, the proposed SEE-MCAM can achieve 8.7x and 4.9x more energy efficiency than ReRAM-based and FeFET-based MCAMs, respectively. Benchmarking results show that our approach provides up to 3 orders of magnitude improvement in speedup and energy efficiency over a GPU implementation in accelerating a novel quantized hyperdimensional computing (HDC) application.

Motivation & Objective

  • To address the memory wall bottleneck in AI workloads by enabling high-density, low-energy associative search within memory.
  • To overcome the limitations of existing binary and multi-bit CAMs based on NVMs, which are constrained by single-level cell (SLC) operation and suffer from high area, energy, or sensitivity to device variation.
  • To exploit the multi-level cell (MLC) characteristics of FeFETs to enable scalable, compact, and robust multi-bit CAM (MCAM) designs with improved energy efficiency and performance.
  • To validate the feasibility and superiority of the proposed FeFET-based MCAM in real-world AI workloads, particularly quantized hyperdimensional computing (HDC).

Proposed method

  • Designing a 2FeFET multi-level cell (MLC) structure using the MIBO (multi-level inverter-based output) concept to store multiple bits per FeFET.
  • Proposing two CAM architectures: NOR-type 2FeFET-1T and NAND-type 2FeFET-2T, to support multi-bit associative search with minimal area and energy overhead.
  • Implementing a non-linear quantization scheme to map hypervector elements to 3-bit values, enabling compatibility with the 3-bit-per-cell SEE-MCAM design.
  • Integrating the SEE-MCAM into a quantized hyperdimensional computing (HDC) framework to benchmark performance and energy efficiency against GPU and other CAM-based implementations.
  • Using a 16T CMOS CAM as baseline for comparison, and evaluating performance via simulation and hardware-aware benchmarks on real datasets (ISOLET, UCIHAR, PAMAP).
  • Employing Nvidia System Management Interface and PyTorch profiler for accurate power and delay measurement in GPU and SEE-MCAM-based HDC workloads.

Experimental results

Research questions

  • RQ1Can FeFET-based multi-bit CAMs achieve higher storage density and energy efficiency than existing binary and multi-bit CAMs?
  • RQ2How can the multi-level cell (MLC) property of FeFETs be effectively exploited to enable scalable, low-energy multi-bit associative search?
  • RQ3What are the performance and energy efficiency gains of SEE-MCAM over CMOS-based CAMs and other NVM-based CAMs in real AI workloads?
  • RQ4To what extent does SEE-MCAM improve accuracy and efficiency in quantized hyperdimensional computing (HDC) applications compared to binary CAMs and GPU implementations?

Key findings

  • The proposed 2FeFET-1T SEE-MCAM achieves 9.8× higher energy efficiency and 1.6× lower search latency compared to 16T CMOS CAM.
  • SEE-MCAM achieves 3 bits per cell functionality, resulting in a 3× improvement in storage density over binary CAMs.
  • The area per bit of the SEE-MCAM cell is only 8% of that of conventional CMOS CAM, demonstrating significant area scaling.
  • Compared to ReRAM-based and FeFET-based MCAMs, SEE-MCAM achieves 8.7× and 4.9× higher energy efficiency, respectively.
  • In quantized HDC workloads, SEE-MCAM-based implementations deliver up to 3 orders of magnitude improvement in both speedup and energy efficiency over GPU (GTX 1080ti) implementations.
  • The 3-bit SEE-MCAM achieves 2.41% higher accuracy than the 2-bit version due to increased hypervector dimensionality, and 2.26% higher accuracy than a binary COSIME-based implementation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.