Cancer classification and pathway discovery using non-negative matrix factorization

09/27/2018
by   Zexian Zeng, et al.
0

Extracting genetic information from a full range of sequencing data is important for understanding diseases. We propose a novel method to effectively explore the landscape of genetic mutations and aggregate them to predict cancer type. We used multinomial logistic regression, nonsmooth non-negative matrix factorization (nsNMF), and support vector machine (SVM) to utilize the full range of sequencing data, aiming at better aggregating genetic mutations and improving their power in predicting cancer types. Specifically, we introduced a classifier to distinguish cancer types using somatic mutations obtained from whole-exome sequencing data. Mutations were identified from multiple cancers and scored using SIFT, PP2, and CADD, and grouped at the individual gene level. The nsNMF was then applied to reduce dimensionality and to obtain coefficient and basis matrices. A feature matrix was derived from the obtained matrices to train a classifier for cancer type classification with the SVM model. We have demonstrated that the classifier was able to distinguish the cancer types with reasonable accuracy. In five-fold cross-validations using mutation counts as features, the average prediction accuracy was 77.1 outperforming baselines and outperforming models using mutation scores as features. Using the factor matrices derived from the nsNMF, we identified multiple genes and pathways that are significantly associated with each cancer type. This study presents a generic and complete pipeline to study the associations between somatic mutations and cancers. The discovered genes and pathways associated with each cancer type can lead to biological insights. The proposed method can be adapted to other studies for disease classification and pathway discovery.

READ FULL TEXT
research
07/15/2020

Prediction of Cancer Microarray and DNA Methylation Data using Non-negative Matrix Factorization

Over the past few years, there has been a considerable spread of microar...
research
08/09/2020

Low-Rank Reorganization via Proportional Hazards Non-negative Matrix Factorization Unveils Survival Associated Gene Clusters

One of the central goals of precision health is the understanding and in...
research
12/01/2017

Bayesian Semi-nonnegative Tri-matrix Factorization to Identify Pathways Associated with Cancer Types

Identifying altered pathways that are associated with specific cancer ty...
research
05/14/2018

Integrating Hypertension Phenotype and Genotype with Hybrid Non-negative Matrix Factorization

Hypertension is a heterogeneous syndrome in need of improved subtyping u...
research
08/30/2023

Multiple Augmented Reduced Rank Regression for Pan-Cancer Analysis

Statistical approaches that successfully combine multiple datasets are m...
research
03/17/2020

Two Tier Prediction of Stroke Using Artificial Neural Networks and Support Vector Machines

Cerebrovascular accident (CVA) or stroke is the rapid loss of brain func...
research
12/15/2020

PANTHER: Pathway Augmented Nonnegative Tensor factorization for HighER-order feature learning

Genetic pathways usually encode molecular mechanisms that can inform tar...

Please sign up or login with your details

Forgot password? Click here to reset