A causal view on compositional data

06/21/2021
by   Elisabeth Ailer, et al.
0

Many scientific datasets are compositional in nature. Important examples include species abundances in ecology, rock compositions in geology, topic compositions in large-scale text corpora, and sequencing count data in molecular biology. Here, we provide a causal view on compositional data in an instrumental variable setting where the composition acts as the cause. Throughout, we pay particular attention to the interpretation of compositional causes from the viewpoint of interventions and crisply articulate potential pitfalls for practitioners. Focusing on modern high-dimensional microbiome sequencing data as a timely illustrative use case, our analysis first reveals that popular one-dimensional information-theoretic summary statistics, such as diversity and richness, may be insufficient for drawing causal conclusions from ecological data. Instead, we advocate for multivariate alternatives using statistical data transformations and regression techniques that take the special structure of the compositional sample space into account. In a comparative analysis on synthetic and semi-synthetic data we show the advantages and limitations of our proposal. We posit that our framework may provide a useful starting point for cause-effect estimation in the context of compositional data.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
01/20/2022

A Guideline for the Statistical Analysis of Compositional Data in Immunology

The study of immune cellular composition is of great scientific interest...
research
09/22/2022

A Bayesian Joint Model for Compositional Mediation Effect Selection in Microbiome Data

Analyzing multivariate count data generated by high-throughput sequencin...
research
01/25/2022

Compositional Cubes: A New Concept for Multi-factorial Compositions

Compositional data are commonly known as multivariate observations carry...
research
12/29/2021

Compositional Data Regression in Insurance with Exponential Family PCA

Compositional data are multivariate observations that carry only relativ...
research
10/24/2021

Compositional data analysis – linear algebra, visualization and interpretation

Compositional data analysis is concerned with multivariate data that hav...
research
07/31/2020

A Compositional Model of Consciousness based on Consciousness-Only

Scientific studies of consciousness rely on objects whose existence is i...
research
01/10/2022

A Statistical Analysis of Compositional Surveys

A common statistical problem is inference from positive-valued multivari...

Please sign up or login with your details

Forgot password? Click here to reset