Recurrent Deep Embedding Networks for Genotype Clustering and Ethnicity Prediction

05/30/2018
by   Md. Rezaul Karim, et al.
0

The understanding of variations in genome sequences assists us in identifying people who are predisposed to common diseases, solving rare diseases, and finding the corresponding population group of the individuals from a larger population group. Although classical machine learning techniques allow researchers to identify groups (i.e. clusters) of related variables, the accuracy, and effectiveness of these methods diminish for large and high-dimensional datasets such as the whole human genome. On the other hand, deep neural network architectures (the core of deep learning) can better exploit large-scale datasets to build complex models. In this paper, we use the K-means clustering approach for scalable genomic data analysis aiming towards clustering genotypic variants at the population scale. Finally, we train a deep belief network (DBN) for predicting the geographic ethnicity. We used the genotype data from the 1000 Genomes Project, which covers the result of genome sequencing for 2504 individuals from 26 different ethnic origins and comprises 84 million variants. Our experimental results, with a focus on accuracy and scalability, show the effectiveness and superiority compared to the state-of-the-art.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/24/2021

Note on the offspring distribution for group testing in the linear regime

The group testing problem is concerned with identifying a small set of k...
research
08/01/2022

Unsupervised machine learning framework for discriminating major variants of concern during COVID-19

Due to the rapid evolution of the SARS-CoV-2 (COVID-19) virus, a number ...
research
10/06/2020

Deep Neural Network: An Efficient and Optimized Machine Learning Paradigm for Reducing Genome Sequencing Error

Genomic data I used in many fields but, it has become known that most of...
research
11/04/2020

A deep learning classifier for local ancestry inference

Local ancestry inference (LAI) identifies the ancestry of each segment o...
research
07/24/2023

Clustering MIC data through Bayesian mixture models: an application to detect M. Tuberculosis resistance mutations

Antimicrobial resistance is becoming a major threat to public health thr...
research
09/12/2021

Spike2Vec: An Efficient and Scalable Embedding Approach for COVID-19 Spike Sequences

With the rapid global spread of COVID-19, more and more data related to ...
research
12/30/2022

Topical Hidden Genome: Discovering Latent Cancer Mutational Topics using a Bayesian Multilevel Context-learning Approach

Statistical inference on the cancer-site specificities of collective ult...

Please sign up or login with your details

Forgot password? Click here to reset