Knowledge distillation for fast and accurate DNA sequence correction

11/17/2022
by   Anastasiya Belyaeva, et al.
0

Accurate genome sequencing can improve our understanding of biology and the genetic basis of disease. The standard approach for generating DNA sequences from PacBio instruments relies on HMM-based models. Here, we introduce Distilled DeepConsensus - a distilled transformer-encoder model for sequence correction, which improves upon the HMM-based methods with runtime constraints in mind. Distilled DeepConsensus is 1.3x faster and 1.5x smaller than its larger counterpart while improving the yield of high quality reads (Q30) over the HMM-based method by 1.69x (vs. 1.73x for larger model). With improved accuracy of genomic sequences, Distilled DeepConsensus improves downstream applications of genomic sequence analysis such as reducing variant calling errors by 39 3.8 Distilled DeepConsensus are similar between faster and slower models.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/20/2023

Embed-Search-Align: DNA Sequence Alignment using Transformer Models

DNA sequence alignment involves assigning short DNA reads to the most pr...
research
10/10/2019

LISA: Towards Learned DNA Sequence Search

Next-generation sequencing (NGS) technologies have enabled affordable se...
research
02/05/2020

FPGA Acceleration of Sequence Alignment: A Survey

Genomics is changing our understanding of humans, evolution, diseases, a...
research
07/11/2023

DNAGPT: A Generalized Pretrained Tool for Multiple DNA Sequence Analysis Tasks

The success of the GPT series proves that GPT can extract general inform...
research
11/24/2021

Deep metric learning improves lab of origin prediction of genetically engineered plasmids

Genome engineering is undergoing unprecedented development and is now be...
research
02/16/2023

ClaPIM: Scalable Sequence CLAssification using Processing-In-Memory

DNA sequence classification is a fundamental task in computational biolo...

Please sign up or login with your details

Forgot password? Click here to reset