Cross-Cluster Weighted Forests

05/17/2021
by   Maya Ramchandran, et al.
8

Adapting machine learning algorithms to better handle the presence of natural clustering or batch effects within training datasets is imperative across a wide variety of biological applications. This article considers the effect of ensembling Random Forest learners trained on clusters within a single dataset with heterogeneity in the distribution of the features. We find that constructing ensembles of forests trained on clusters determined by algorithms such as k-means results in significant improvements in accuracy and generalizability over the traditional Random Forest algorithm. We denote our novel approach as the Cross-Cluster Weighted Forest, and examine its robustness to various data-generating scenarios and outcome models. Furthermore, we explore the influence of the data-partitioning and ensemble weighting strategies on conferring the benefits of our method over the existing paradigm. Finally, we apply our approach to cancer molecular profiling and gene expression datasets that are naturally divisible into clusters and illustrate that our approach outperforms classic Random Forest. Code and supplementary material are available at https://github.com/m-ramchandran/cross-cluster.

READ FULL TEXT
research
09/01/2020

Improved Weighted Random Forest for Classification Problems

Several studies have shown that combining machine learning models in an ...
research
05/02/2014

Asymptotic Theory for Random Forests

Random forests have proven to be reliable predictive algorithms in many ...
research
10/26/2020

Data Segmentation via t-SNE, DBSCAN, and Random Forest

This research proposes a data segmentation technique which is easy to in...
research
12/10/2022

Scaling pattern mining through non-overlapping variable partitioning

Biclustering algorithms play a central role in the biotechnological and ...
research
02/17/2021

BEDS: Bagging ensemble deep segmentation for nucleus segmentation with testing stage stain augmentation

Reducing outcome variance is an essential task in deep learning based me...
research
01/05/2023

Random forests, sound symbolism and Pokemon evolution

This study constructs machine learning algorithms that are trained to cl...
research
04/05/2020

An Unsupervised Random Forest Clustering Technique for Automatic Traffic Scenario Categorization

A modification of the Random Forest algorithm for the categorization of ...

Please sign up or login with your details

Forgot password? Click here to reset