[Paper Review] Achievable Rates of Concatenated Codes in DNA Storage under Substitution Errors
This paper investigates concatenated coding schemes for DNA storage under substitution errors, showing that standard schemes with one strand per inner block fail to achieve capacity due to unordered, randomly accessed strands. By grouping multiple strands into a single inner block, the authors propose a modified scheme that significantly narrows the gap to capacity, achieving rates close to the theoretical limit through Monte Carlo simulations and asymptotic analysis.
In this paper, we study achievable rates of concatenated coding schemes over a deoxyribonucleic acid (DNA) storage channel. Our channel model incorporates the main features of DNA-based data storage. First, information is stored on many, short DNA strands. Second, the strands are stored in an unordered fashion inside the storage medium and each strand is replicated many times. Third, the data is accessed in an uncontrollable manner, i.e., random strands are drawn from the medium and received, possibly with errors. As one of our results, we show that there is a significant gap between the channel capacity and the achievable rate of a standard concatenated code in which one strand corresponds to an inner block. This is in fact surprising as for other channels, such as $q$-ary symmetric channels, concatenated codes are known to achieve the capacity. We further propose a modified concatenated coding scheme by combining several strands into one inner block, which allows to narrow the gap and achieve rates that are close to the capacity.
Motivation & Objective
- To analyze the achievable rates of standard concatenated codes in DNA storage under unordered, randomly accessed strands with substitution errors.
- To identify why standard concatenated codes fail to achieve capacity despite success on other channels.
- To design an improved coding scheme that approaches the channel capacity by redefining the inner block structure.
- To provide a practical, efficiently encodable and decodable coding solution for real-world DNA storage systems.
Proposed method
- Model the DNA storage channel as a probabilistic channel with unordered input and output strands, where each strand is subject to substitution errors.
- Define a modified concatenated coding scheme where multiple DNA strands are grouped into a single inner block, increasing resilience to unordered access.
- Use a typicality-like decoder with random coding arguments to derive achievable rates, leveraging Poisson-distributed strand draws.
- Approximate the achievable rate via Monte Carlo simulation by sampling from a K-variate Poisson distribution with parameter c.
- Analyze the asymptotic behavior as block length K → ∞, showing convergence to capacity under the new scheme.
- Derive the overall rate R = Rin × Rout × (1 − β/Rix) and optimize Rin to maximize R for finite K.
Experimental results
Research questions
- RQ1Why do standard concatenated codes with one strand per inner block fail to achieve capacity in DNA storage despite their success on other channels?
- RQ2What is the fundamental reason for the gap between channel capacity and achievable rate in DNA storage under unordered strand access?
- RQ3Can redefining the inner block to include multiple strands close the gap to capacity in DNA storage?
- RQ4How does the achievable rate scale with block length K under the modified coding scheme?
- RQ5What is the optimal choice of inner code rate Rin to maximize the overall transmission rate R for finite K?
Key findings
- The standard concatenated code with one strand per inner block cannot achieve capacity due to the unordered, randomly accessed nature of DNA storage, even with optimal inner and outer codes.
- The proposed modified scheme, where multiple strands form a single inner block, achieves rates significantly closer to the channel capacity than the standard approach.
- For K = 100, the achievable rate is substantially higher than for K = 1, demonstrating that increasing block length closes the gap to capacity.
- The asymptotic achievable rate as K → ∞ approaches the channel capacity, confirming the scheme's capacity-approaching potential.
- Monte Carlo simulations confirm that the overall rate R approaches the theoretical asymptotic limit for growing K, with convergence observed at K ≥ 100.
- The optimal inner code rate Rin is found to maximize R, and its position shifts with K, providing a design criterion for practical implementation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.