Selective Inference for Hierarchical Clustering

12/05/2020
by   Lucy L. Gao, et al.
0

Testing for a difference in means between two groups is fundamental to answering research questions across virtually every scientific area. Classical tests control the Type I error rate when the groups are defined a priori. However, when the groups are instead defined via a clustering algorithm, then applying a classical test for a difference in means between the groups yields an extremely inflated Type I error rate. Notably, this problem persists even if two separate and independent data sets are used to define the groups and to test for a difference in their means. To address this problem, in this paper, we propose a selective inference approach to test for a difference in means between two clusters obtained from any clustering method. Our procedure controls the selective Type I error rate by accounting for the fact that the null hypothesis was generated from the data. We describe how to efficiently compute exact p-values for clusters obtained using agglomerative hierarchical clustering with many commonly used linkages. We apply our method to simulated data and to single-cell RNA-seq data.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/29/2022

Selective inference for k-means clustering

We consider the problem of testing for a difference in means between clu...
research
10/24/2022

Post-clustering difference testing: valid inference and practical considerations

Clustering is part of unsupervised analysis methods that consist in grou...
research
06/15/2021

Tree-Values: selective inference for regression trees

We consider conducting inference on the output of the Classification and...
research
09/04/2023

Selective inference after convex clustering with ℓ_1 penalization

Classical inference methods notoriously fail when applied to data-driven...
research
03/14/2021

Quantifying uncertainty in spikes estimated from calcium imaging data

In recent years, a number of methods have been proposed to estimate the ...
research
10/17/2018

Structural Equation Modeling and simultaneous clustering through the Partial Least Squares algorithm

The identification of different homogeneous groups of observations and t...
research
09/21/2021

More powerful selective inference for the graph fused lasso

The graph fused lasso – which includes as a special case the one-dimensi...

Please sign up or login with your details

Forgot password? Click here to reset