[Paper Review] A High-Throughput Multi-Mode LDPC Decoder for 5G NR
This paper presents a high-throughput, multi-mode LDPC decoder for 5G NR using a partially parallel architecture with a flooding schedule, supporting code rates and lengths up to a lifting size of 96. It dynamically increases parallelism for smaller lifting sizes (Z ≤ 48 and Z ≤ 24) to maintain peak throughput, achieving 13.46 Gbps in 28 nm CMOS with 1.03 mm² area and 229 mW power consumption.
This paper presents a partially parallel low-density parity-check (LDPC) decoder designed for the 5G New Radio (NR) standard. The design is using a multi-block parallel architecture with a flooding schedule. The decoder can support any code rates and code lengths up to the lifting size Zmax= 96. To compensate for the dropped throughput associated with the smaller Z values, the design can double and quadruple its parallelism when lifting sizes Z<= 48 and Z<= 24 are selected respectively. Therefore, the decoder can process up to eight frames and restore the throughput to the maximum. To simplify the design's architecture, a new variable node for decoding the extended parity bits present in the lower code rates is proposed. The FPGA implementation of the decoder results in a throughput of 2.1 Gbps decoding the 11/12 code rate. Additionally, the synthesized decoder using the 28 nm TSMC technology, achieves a maximum clock frequency of 526 MHz and a throughput of 13.46 Gbps. The core decoder occupies 1.03 mm2, and the power consumption is 229 mW.
Motivation & Objective
- Address the need for high-throughput, flexible LDPC decoding in 5G New Radio (NR) systems.
- Support variable code rates and code lengths up to a maximum lifting size of 96.
- Maintain high throughput across all code rates, especially for lower rates where throughput typically drops.
- Simplify the decoder architecture by introducing a novel variable node for handling extended parity bits in low-rate codes.
- Achieve high performance in both FPGA and ASIC implementations with minimal area and power overhead.
Proposed method
- Adopt a partially parallel architecture with a flooding schedule to balance area and throughput.
- Implement dynamic parallelism: double parallelism for Z ≤ 48 and quadruple for Z ≤ 24 to compensate for reduced throughput at smaller lifting sizes.
- Introduce a new variable node structure to efficiently decode extended parity bits in low-rate codes, reducing architectural complexity.
- Use a multi-block parallel design to distribute computation across multiple processing units and improve throughput scaling.
- Optimize the decoder for 28 nm TSMC technology via synthesis, achieving high clock frequency and low power consumption.
- Validate performance via FPGA prototyping and ASIC synthesis, focusing on throughput, area, and power metrics.
Experimental results
Research questions
- RQ1How can an LDPC decoder maintain high throughput across all 5G NR code rates, especially for low-rate codes with extended parity bits?
- RQ2What architectural techniques can dynamically scale parallelism to compensate for throughput degradation at smaller lifting sizes?
- RQ3How can the complexity of handling extended parity bits in low-rate codes be reduced without sacrificing decoding performance?
- RQ4What is the achievable throughput, area, and power efficiency of a high-throughput LDPC decoder in 28 nm CMOS technology?
- RQ5Can a single multi-mode decoder design effectively support all 5G NR code rates and lengths with minimal area and power overhead?
Key findings
- The FPGA implementation achieves a throughput of 2.1 Gbps when decoding the 11/12 code rate.
- The ASIC synthesis in 28 nm TSMC technology achieves a maximum clock frequency of 526 MHz and a peak throughput of 13.46 Gbps.
- The core decoder occupies only 1.03 mm² of area, demonstrating high area efficiency.
- Power consumption is measured at 229 mW under peak operation, indicating low power efficiency for high-throughput operation.
- The dynamic parallelism mechanism successfully restores peak throughput for smaller lifting sizes (Z ≤ 48 and Z ≤ 24) by doubling and quadrupling parallelism respectively.
- The proposed variable node for extended parity bits simplifies the architecture and reduces hardware complexity in low-rate decoding modes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.