Kernel learning approaches for summarising and combining posterior similarity matrices

09/27/2020
by   Alessandra Cabassi, et al.
12

When using Markov chain Monte Carlo (MCMC) algorithms to perform inference for Bayesian clustering models, such as mixture models, the output is typically a sample of clusterings (partitions) drawn from the posterior distribution. In practice, a key challenge is how to summarise this output. Here we build upon the notion of the posterior similarity matrix (PSM) in order to suggest new approaches for summarising the output of MCMC algorithms for Bayesian clustering models. A key contribution of our work is the observation that PSMs are positive semi-definite, and hence can be used to define probabilistically-motivated kernel matrices that capture the clustering structure present in the data. This observation enables us to employ a range of kernel methods to obtain summary clusterings, and otherwise exploit the information summarised by PSMs. For example, if we have multiple PSMs, each corresponding to a different dataset on a common set of statistical units, we may use standard methods for combining kernels in order to perform integrative clustering. We may moreover embed PSMs within predictive kernel models in order to perform outcome-guided data integration. We demonstrate the performances of the proposed methods through a range of simulation studies as well as two real data applications. R code is available at https://github.com/acabassi/combine-psms.

READ FULL TEXT

page 22

page 29

page 31

page 36

page 38

page 39

page 40

page 41

research
12/20/2021

Bayesian nonparametric model based clustering with intractable distributions: an ABC approach

Bayesian nonparametric mixture models offer a rich framework for model b...
research
03/31/2023

Bayesian Clustering via Fusing of Localized Densities

Bayesian clustering typically relies on mixture models, with each compon...
research
04/15/2019

Multiple kernel learning for integrative consensus clustering of genomic datasets

Diverse applications - particularly in tumour subtyping - have demonstra...
research
05/17/2022

BayesMix: Bayesian Mixture Models in C++

We describe BayesMix, a C++ library for MCMC posterior simulation for ge...
research
02/15/2019

BAREB: A Bayesian repulsive biclustering model for periodontal data

Preventing periodontal diseases (PD) and maintaining the structure and f...
research
04/10/2023

Scalable Randomized Kernel Methods for Multiview Data Integration and Prediction

We develop scalable randomized kernel methods for jointly associating da...
research
03/31/2021

pivmet: Pivotal Methods for Bayesian Relabelling and k-Means Clustering

The identification of groups' prototypes, i.e. elements of a dataset tha...

Please sign up or login with your details

Forgot password? Click here to reset