PIntMF: Penalized Integrative Matrix Factorization Method for Multi-Omics Data

03/03/2021
by   Morgane Pierre-Jean, et al.
0

It is more and more common to explore the genome at diverse levels and not only at a single omic level. Through integrative statistical methods, omics data have the power to reveal new biological processes, potential biomarkers, and subgroups of a cohort. The matrix factorization (MF) is a unsupervised statistical method that allows giving a clustering of individuals, but also revealing relevant omic variables from the various blocks. Here, we present PIntMF (Penalized Integrative Matrix Factorization), a model of MF with sparsity, positivity and equality constraints.To induce sparsity in the model, we use a classical Lasso penalization on variable and individual matrices. For the matrix of samples, sparsity helps for the clustering, and normalization (matching an equality constraint) of inferred coefficients is added for a better interpretation. Besides, we add an automatic tuning of the sparsity parameters using the famous glmnet package. We also proposed three criteria to help the user to choose the number of latent variables. PIntMF was compared to other state-of-the-art integrative methods including feature selection techniques in both synthetic and real data. PIntMF succeeds in finding relevant clusters as well as variables in two types of simulated data (correlated and uncorrelated). Then, PIntMF was applied to two real datasets (Diet and cancer), and it reveals interpretable clusters linked to available clinical data. Our method outperforms the existing ones on two criteria (clustering and variable selection). We show that PIntMF is an easy, fast, and powerful tool to extract patterns and cluster samples from multi-omics data.

READ FULL TEXT

page 7

page 10

page 11

research
04/27/2021

Structured Sparse Non-negative Matrix Factorization with L20-Norm for scRNA-seq Data Analysis

Non-negative matrix factorization (NMF) is a powerful tool for dimension...
research
05/25/2023

Flexible Variable Selection for Clustering and Classification

The importance of variable selection for clustering has been recognized ...
research
06/03/2019

Clustering by Orthogonal NMF Model and Non-Convex Penalty Optimization

The non-negative matrix factorization (NMF) model with an additional ort...
research
01/29/2015

Tensor Factorization via Matrix Factorization

Tensor factorization arises in many machine learning applications, such ...
research
05/17/2020

Bayesian biclustering for microbial metagenomic sequencing data via multinomial matrix factorization

High-throughput sequencing technology provides unprecedented opportuniti...
research
06/09/2019

Integrative Factorization of Bidimensionally Linked Matrices

Advances in molecular "omics'" technologies have motivated new methodolo...

Please sign up or login with your details

Forgot password? Click here to reset