Return
A Segmented-Edit Error-Correcting Code With Re-Synchronization Function for DNA-Based Storage Systems
DOI:10.1109/TETC.2022.3225570.png)
Abstract
En 中文
As a powerful tool for storing digital information in chemically synthesized molecules, DNA based data storage has undergone continuous development and received increasingly more attention. Effi- ciently recovering information from large-scale DNA strands that suffer from insertions, deletions, and substitution errors (collectively referred to as edit errors), is one of the major bottlenecks in DNA-based storage systems. To cope with this challenge, in this paper, we provide a segmented-edit error-correcting code with the re-synchronization function, termed the DNA-LM code. Compared with the previous segmented-error correcting codes, it has a systematic structure and does not require the endpoint of the received segment as pre-requisite information for decoding. In the case that the number of edit errors exceeds the edit error -correcting capability of a segment, it can easily regain synchronization to ensure that the subsequent decoding continues. Both encoding and decoding complexity is linear in the codeword length. The redundancy of each segment is | log k | + 6 quaternary symbols, where k is the length of the message segment. We further generalize the decoding algorithm to deal with duplicated DNA strands, whereas it still maintains linear time complexity in the codeword length and the number of duplications. Simulations under a stochastic edit errors model show that, at a low raw error rate of the next-gen sequencing, our code can enable error-free decoding by concatenating with the (255,223) RS code.
Keywords:
DNA-based storage system
segmented-edit error-correcting codes
synchronization errors
VT codes
Journal
IF:
5.4
Papers:
1.1K
Citations:
3.4K

