Statistical Depth based Normalization and Outlier Detection of Gene Expression Data

06/28/2022
by   Alicia Nieto-Reyes, et al.
0

Normalization and outlier detection belong to the preprocessing of gene expression data. We propose a natural normalization procedure based on statistical data depth which normalizes to the distribution of gene expressions of the most representative gene expression of the group. This differ from the standard method of quantile normalization, based on the coordinate-wise median array that lacks of the well-known properties of the one-dimensional median. The statistical data depth maintains those good properties. Gene expression data are known for containing outliers. Although detecting outlier genes in a given gene expression dataset has been broadly studied, these methodologies do not apply for detecting outlier samples, given the difficulties posed by the high dimensionality but low sample size structure of the data. The standard procedures used for detecting outlier samples are visual and based on dimension reduction techniques; instances are multidimensional scaling and spectral map plots. For detecting outlier genes in a given gene expression dataset, we propose an analytical procedure and based on the Tukey's concept of outlier and the notion of statistical depth, as previous methodologies lead to unassertive and wrongful outliers. We reveal the outliers of four datasets; as a necessary step for further research.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/12/2021

Autoencoding Under Normalization Constraints

Likelihood is a standard estimate for outlier detection. The specific ro...
research
12/13/2017

Multiple testing for outlier detection in functional data

We propose a novel procedure for outlier detection in functional data, i...
research
03/29/2020

The covariance shift (C-SHIFT) algorithm for normalizing biological data

Omics technologies are powerful tools for analyzing patterns in gene exp...
research
12/19/2018

Covariance-based sample selection for heterogenous data: Applications to gene expression and autism risk gene detection

Risk for autism can be influenced by genetic mutations in hundreds of ge...
research
06/29/2015

Integrative analysis of gene expression and phenotype data

The linking genotype to phenotype is the fundamental aim of modern genet...
research
04/19/2021

Multidimensional Scaling for Gene Sequence Data with Autoencoders

Multidimensional scaling of gene sequence data has long played a vital r...
research
03/01/2018

Modeling Data Containing Outliers using ARIMA Additive Outlier (ARIMA-AO)

The aim this study is discussed on the detection and correction of data ...

Please sign up or login with your details

Forgot password? Click here to reset