Insertion and Deletion Correction in Polymer-based Data Storage

01/21/2022
by   Anisha Banerjee, et al.
0

Synthetic polymer-based storage seems to be a particularly promising candidate that could help to cope with the ever-increasing demand for archival storage requirements. It involves designing molecules of distinct masses to represent the respective bits {0,1}, followed by the synthesis of a polymer of molecular units that reflects the order of bits in the information string. Reading out the stored data requires the use of a tandem mass spectrometer, that fragments the polymer into shorter substrings and provides their corresponding masses, from which the composition, i.e. the number of 1s and 0s in the concerned substring can be inferred. Prior works have dealt with the problem of unique string reconstruction from the set of all possible compositions, called composition multiset. This was accomplished either by determining which string lengths always allow unique reconstruction, or by formulating coding constraints to facilitate the same for all string lengths. Additionally, error-correcting schemes to deal with substitution errors caused by imprecise fragmentation during the readout process, have also been suggested. This work builds on this research by generalizing previously considered error models, mainly confined to substitution of compositions. To this end, we define new error models that consider insertions of spurious compositions and deletions of existing ones, thereby corrupting the composition multiset. We analyze if the reconstruction codebook proposed by Pattabiraman et al. is indeed robust to such errors, and if not, propose new coding constraints to remedy this.

READ FULL TEXT
research
01/14/2020

Mass Error-Correction Codes for Polymer-Based Data Storage

We consider the problem of correcting mass readout errors in information...
research
04/19/2019

Reconstruction and Error-Correction Codes for Polymer-Based Data Storage

Motivated by polymer-based data-storage platforms that use chains of bin...
research
03/02/2020

Coding for Polymer-Based Data Storage

Motivated by polymer-based data-storage platforms that use chains of bin...
research
04/12/2018

Unique Reconstruction of Coded Strings from Multiset Substring Spectra

The problem of reconstructing strings from their substring spectra has a...
research
01/24/2022

A New Algebraic Approach for String Reconstruction from Substring Compositions

We consider the problem of binary string reconstruction from the multise...
research
05/17/2023

Error-Correcting Codes for Nanopore Sequencing

Nanopore sequencers, being superior to other sequencing technologies for...

Please sign up or login with your details

Forgot password? Click here to reset