Multiple kernel learning for integrative consensus clustering of genomic datasets

04/15/2019
by   Alessandra Cabassi, et al.
0

Diverse applications - particularly in tumour subtyping - have demonstrated the importance of integrative clustering as a means to combine information from multiple high-dimensional omics datasets. Cluster-Of-Clusters Analysis (COCA) is a popular integrative clustering method that has been widely applied in the context of tumour subtyping. However, the properties of COCA have never been systematically explored, and the robustness of this approach to the inclusion of noisy datasets, or datasets that define conflicting clustering structures, is unclear. We rigorously benchmark COCA, and present Kernel Learning Integrative Clustering (KLIC) as an alternative strategy. KLIC frames the challenge of combining clustering structures as a multiple kernel learning problem, in which different datasets each provide a weighted contribution to the final clustering. This allows the contribution of noisy datasets to be down-weighted relative to more informative datasets. We show through extensive simulation studies that KLIC is more robust than COCA in a variety of situations. R code to run KLIC and COCA can be found at https://github.com/acabassi/klic

READ FULL TEXT
research
07/05/2022

Local Sample-weighted Multiple Kernel Clustering with Consensus Discriminative Graph

Multiple kernel clustering (MKC) is committed to achieving optimal infor...
research
09/27/2020

Kernel learning approaches for summarising and combining posterior similarity matrices

When using Markov chain Monte Carlo (MCMC) algorithms to perform inferen...
research
11/02/2021

The climatic interdependence of extreme-rainfall events around the globe

The identification of regions of similar climatological behavior can be ...
research
09/19/2010

Pair-Wise Cluster Analysis

This paper studies the problem of learning clusters which are consistent...
research
06/27/2018

Quantile-based clustering

A new cluster analysis method, K-quantiles clustering, is introduced. K-...
research
12/04/2014

Iterative Subsampling in Solution Path Clustering of Noisy Big Data

We develop an iterative subsampling approach to improve the computationa...
research
09/26/2017

Scale Adaptive Clustering of Multiple Structures

We propose the segmentation of noisy datasets into Multiple Inlier Struc...

Please sign up or login with your details

Forgot password? Click here to reset