Repeated undersampling in PrInDT (RePrInDT): Variation in undersampling and prediction, and ranking of predictors in ensembles

08/11/2021
by   Claus Weihs, et al.
0

In this paper, we extend our PrInDT method (Weihs Buschfeld 2021a) towards undersampling with different percentages of the smaller and the larger classes (psmall and plarge), stratification of predictors, varying the prediction threshold, and measuring variable importance in ensembles. An application of these methods to a linguistic example suggests the following: 1. In undersampling, a careful selection of the percentages plarge and psmall is important for building models with high balanced accuracies; 2. Stratification of predictors does not majorly enhance balanced accuracies; 3. Lowering the prediction threshold for the smaller class turns out to be an alternative method to undersampling because it increases the likelihood of the smaller class being selected. Finally, we introduce a method for ranking predictor importance that allows for a straightforward interpretation of the results.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/27/2021

NesPrInDT: Nested undersampling in PrInDT

In this paper, we extend our PrInDT method (Weihs, Buschfeld 2021) towar...
research
08/01/2022

Accelerated and interpretable oblique random survival forests

The oblique random survival forest (RSF) is an ensemble supervised learn...
research
03/03/2021

Combining Prediction and Interpretation in Decision Trees (PrInDT) – a Linguistic Example

In this paper, we show that conditional inference trees and ensembles ar...
research
04/24/2023

A Cheat Sheet for Bayesian Prediction

This paper reviews the growing field of Bayesian prediction. Bayes point...
research
03/30/2023

KOO approach for scalable variable selection problem in large-dimensional regression

An important issue in many multivariate regression problems is to elimin...
research
06/28/2016

Reviving Threshold-Moving: a Simple Plug-in Bagging Ensemble for Binary and Multiclass Imbalanced Data

Class imbalance presents a major hurdle in the application of data minin...
research
03/15/2022

Approximability and Generalisation

Approximate learning machines have become popular in the era of small de...

Please sign up or login with your details

Forgot password? Click here to reset