Robust Regression with Compositional Covariates

by   Aditya Mishra, et al.

Many high-throughput sequencing data sets in biology are compositional in nature. A prominent example is microbiome profiling data, including targeted amplicon-based and metagenomic sequencing data. These profiling data comprises surveys of microbial communities in their natural habitat and sparse proportional (or compositional) read counts that represent operational taxonomic units or genes. When paired measurements of other covariates, including physicochemical properties of the habitat or phenotypic variables of the host, are available, inference of parsimonious and robust statistical relationships between the microbial abundance data and the covariate measurements is often an important first step in exploratory data analysis. To this end, we propose a sparse robust statistical regression framework that considers compositional and non-compositional measurements as predictors and identifies outliers in continuous response variables. Our model extends the seminal log-contrast model of Aitchison and Bacon-Shone (1984) by a mean shift formulation for capturing outliers, sparsity-promoting convex and non-convex penalties for parsimonious model selection, and data-driven robust initialization procedures adapted to the compositional setting. We show, in theory and simulations, the ability of our approach to jointly select a sparse set of predictive microbial features and identify outliers in the response. We illustrate the viability of our method by robustly predicting human body mass indices from American Gut Project amplicon data and non-compositional covariate data. We believe that the robust estimators introduced here and available in the R package RobRegCC can serve as a practical tool for reliable statistical regression analysis of compositional data, including microbiome survey data.



There are no comments yet.


page 1

page 2

page 3

page 4


Robust Differential Abundance Test in Compositional Data

Differential abundance tests in the compositional data are essential and...

Supervised Learning and Model Analysis with Compositional Data

The compositionality and sparsity of high-throughput sequencing data pos...

High-dimensional Log-Error-in-Variable Regression with Applications to Microbial Compositional Data Analysis

In microbiome and genomic study, the regression of compositional data ha...

Testing for differential abundance in compositional counts data, with application to microbiome studies

In order to identify which taxa differ in the microbiome community acros...

Regularized regression on compositional trees with application to MRI analysis

A compositional tree refers to a tree structure on a set of random varia...

A Statistical Analysis of Compositional Surveys

A common statistical problem is inference from positive-valued multivari...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.