Multiple Augmented Reduced Rank Regression for Pan-Cancer Analysis

08/30/2023
by   Jiuzhou Wang, et al.
0

Statistical approaches that successfully combine multiple datasets are more powerful, efficient, and scientifically informative than separate analyses. To address variation architectures correctly and comprehensively for high-dimensional data across multiple sample sets (i.e., cohorts), we propose multiple augmented reduced rank regression (maRRR), a flexible matrix regression and factorization method to concurrently learn both covariate-driven and auxiliary structured variation. We consider a structured nuclear norm objective that is motivated by random matrix theory, in which the regression or factorization terms may be shared or specific to any number of cohorts. Our framework subsumes several existing methods, such as reduced rank regression and unsupervised multi-matrix factorization approaches, and includes a promising novel approach to regression and factorization of a single dataset (aRRR) as a special case. Simulations demonstrate substantial gains in power from combining multiple datasets, and from parsimoniously accounting for all structured variation. We apply maRRR to gene expression data from multiple cancer types (i.e., pan-cancer) from TCGA, with somatic mutations as covariates. The method performs well with respect to prediction and imputation of held-out data, and provides new insights into mutation-driven and auxiliary variation that is shared or specific to certain cancer types.

READ FULL TEXT

page 4

page 32

page 34

page 37

research
02/07/2020

Bidimensional linked matrix factorization for pan-omics pan-cancer analysis

Several modern applications require the integration of multiple large da...
research
06/09/2019

Integrative Factorization of Bidimensionally Linked Matrices

Advances in molecular "omics'" technologies have motivated new methodolo...
research
01/09/2013

Nonparametric Reduced Rank Regression

We propose an approach to multivariate nonparametric regression that gen...
research
08/09/2020

Low-Rank Reorganization via Proportional Hazards Non-negative Matrix Factorization Unveils Survival Associated Gene Clusters

One of the central goals of precision health is the understanding and in...
research
09/27/2018

Cancer classification and pathway discovery using non-negative matrix factorization

Extracting genetic information from a full range of sequencing data is i...
research
11/29/2022

Bayesian Simultaneous Factorization and Prediction Using Multi-Omic Data

Understanding of the pathophysiology of obstructive lung disease (OLD) i...
research
04/30/2019

Composite local low-rank structure in learning drug sensitivity

The molecular characterization of tumor samples by multiple omics data s...

Please sign up or login with your details

Forgot password? Click here to reset