Probabilistic Random Forest: A machine learning algorithm for noisy datasets

11/14/2018
by   Itamar Reis, et al.
0

Machine learning (ML) algorithms become increasingly important in the analysis of astronomical data. However, since most ML algorithms are not designed to take data uncertainties into account, ML based studies are mostly restricted to data with high signal-to-noise ratio. Astronomical datasets of such high-quality are uncommon. In this work we modify the long-established Random Forest (RF) algorithm to take into account uncertainties in the measurements (i.e., features) as well as in the assigned classes (i.e., labels). To do so, the Probabilistic Random Forest (PRF) algorithm treats the features and labels as probability distribution functions, rather than deterministic quantities. We perform a variety of experiments where we inject different types of noise to a dataset, and compare the accuracy of the PRF to that of RF. The PRF outperforms RF in all cases, with a moderate increase in running time. We find an improvement in classification accuracy of up to 10 the case of noisy features, and up to 30 accuracy decreased by less then 5 misclassified objects, compared to a clean dataset. Apart from improving the prediction accuracy in noisy datasets, the PRF naturally copes with missing values in the data, and outperforms RF when applied to a dataset with different noise characteristics in the training and test sets, suggesting that it can be used for Transfer Learning.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
04/19/2018

A Dynamic Boosted Ensemble Learning Based on Random Forest

We propose Dynamic Boosted Random Forest (DBRF), a novel ensemble algori...
research
02/25/2022

MUC-driven Feature Importance Measurement and Adversarial Analysis for Random Forest

The broad adoption of Machine Learning (ML) in security-critical fields ...
research
12/10/2020

A machine learning approach to galaxy properties: Joint redshift - stellar mass probability distributions with Random Forest

We demonstrate that highly accurate joint redshift - stellar mass PDFs c...
research
10/01/2018

Using Machine Learning to Discern Eruption in Noisy Environments: A Case Study using CO2-driven Cold-Water Geyser in Chimayo, New Mexico

We present an approach based on machine learning (ML) to distinguish eru...
research
12/13/2021

Incorporating Measurement Error in Astronomical Object Classification

Most general-purpose classification methods, such as support-vector mach...
research
01/14/2022

Machine Learning of polymer types from the spectral signature of Raman spectroscopy microplastics data

The tools and technology that are currently used to analyze chemical com...
research
07/27/2021

Spatial prediction of apartment rent using regression-based and machine learning-based approaches with a large dataset

Employing a large dataset (at most, the order of n = 10^6), this study a...

Please sign up or login with your details

Forgot password? Click here to reset