Cell Identity Codes: Understanding Cell Identity from Gene Expression Profiles using Deep Neural Networks

06/13/2018
by   Farzad Abdolhosseini, et al.
1

Understanding cell identity is an important task in many biomedical areas. Expression patterns of specific marker genes have been used to characterize some limited cell types, but exclusive markers are not available for many cell types. A second approach is to use machine learning to discriminate cell types based on the whole gene expression profiles (GEPs). The accuracies of simple classification algorithms such as linear discriminators or support vector machines are limited due to the complexity of biological systems. We used deep neural networks to analyze 1040 GEPs from 16 different human tissues and cell types. After comparing different architectures, we identified a specific structure of deep autoencoders that can encode a GEP into a vector of 30 numeric values, which we call the cell identity code (CIC). The original GEP can be reproduced from the CIC with an accuracy comparable to technical replicates of the same experiment. Although we use an unsupervised approach to train the autoencoder, we show different values of the CIC are connected to different biological aspects of the cell, such as different pathways or biological processes. This network can use CIC to reproduce the GEP of the cell types it has never seen during the training. It also can resist some noise in the measurement of the GEP. Furthermore, we introduce classifier autoencoder, an architecture that can accurately identify cell type based on the GEP or the CIC.

READ FULL TEXT

page 16

page 18

page 19

page 20

research
02/04/2016

Discovering Neuronal Cell Types and Their Gene Expression Profiles Using a Spatial Point Process Mixture Model

Cataloging the neuronal cell types that comprise circuitry of individual...
research
06/06/2020

Extracting Cellular Location of Human Proteins Using Deep Learning

Understanding and extracting the patterns of microscopy images has been ...
research
12/09/2019

Modeling somatic computation with non-neural bioelectric networks

The field of basal cognition seeks to understand how adaptive, context-s...
research
10/25/2022

A single-cell gene expression language model

Gene regulation is a dynamic process that connects genotype and phenotyp...
research
09/17/2020

Identification of Biomarkers Controlling Cell Fate In Blood Cell Development

A blood cell lineage consists of several consecutive developmental stage...
research
06/15/2021

Active feature selection discovers minimal gene-sets for classifying cell-types and disease states in single-cell mRNA-seq data

Sequencing costs currently prohibit the application of single cell mRNA-...
research
09/30/2022

Disentangling with Biological Constraints: A Theory of Functional Cell Types

Neurons in the brain are often finely tuned for specific task variables....

Please sign up or login with your details

Forgot password? Click here to reset