Single-Read Reconstruction for DNA Data Storage Using Transformers

09/12/2021
by   Yotam Nahum, et al.
0

As the global need for large-scale data storage is rising exponentially, existing storage technologies are approaching their theoretical and functional limits in terms of density and energy consumption, making DNA based storage a potential solution for the future of data storage. Several studies introduced DNA based storage systems with high information density (petabytes/gram). However, DNA synthesis and sequencing technologies yield erroneous outputs. Algorithmic approaches for correcting these errors depend on reading multiple copies of each sequence and result in excessive reading costs. The unprecedented success of Transformers as a deep learning architecture for language modeling has led to its repurposing for solving a variety of tasks across various domains. In this work, we propose a novel approach for single-read reconstruction using an encoder-decoder Transformer architecture for DNA based data storage. We address the error correction process as a self-supervised sequence-to-sequence task and use synthetic noise injection to train the model using only the decoded reads. Our approach exploits the inherent redundancy of each decoded file to learn its underlying structure. To demonstrate our proposed approach, we encode text, image and code-script files to DNA, produce errors with high-fidelity error simulator, and reconstruct the original files from the noisy reads. Our model achieves lower error rates when reconstructing the original data from a single read of each DNA strand compared to state-of-the-art algorithms using 2-3 copies. This is the first demonstration of using deep learning models for single-read reconstruction in DNA based storage which allows for the reduction of the overall cost of the process. We show that this approach is applicable for various domains and can be generalized to new domains as well.

READ FULL TEXT
research
10/20/2022

Robust Multi-Read Reconstruction from Contaminated Clusters Using Deep Neural Network for DNA Storage

DNA has immense potential as an emerging data storage medium. The princi...
research
05/11/2022

DNA data storage, sequencing data-carrying DNA

DNA is a leading candidate as the next archival storage media due to its...
research
08/31/2021

Deep DNA Storage: Scalable and Robust DNA Storage via Coding Theory and Deep Learning

The concept of DNA storage was first suggested in 1959 by Richard Feynma...
research
02/19/2021

Efficient approximation of DNA hybridisation using deep learning

Deoxyribonucleic acid (DNA) has shown great promise in enabling computat...
research
09/28/2019

Deep Multiple Instance Learning for Taxonomic Classification of Metagenomic read sets

Metagenomic studies have increasingly utilized sequencing technologies i...
research
10/22/2019

Image processing in DNA

The main obstacles for the practical deployment of DNA-based data storag...
research
09/04/2023

Blind Biological Sequence Denoising with Self-Supervised Set Learning

Biological sequence analysis relies on the ability to denoise the imprec...

Please sign up or login with your details

Forgot password? Click here to reset