Feature Ranking for Semi-supervised Learning

08/10/2020
by   Matej Petković, et al.
0

The data made available for analysis are becoming more and more complex along several directions: high dimensionality, number of examples and the amount of labels per example. This poses a variety of challenges for the existing machine learning methods: coping with dataset with a large number of examples that are described in a high-dimensional space and not all examples have labels provided. For example, when investigating the toxicity of chemical compounds there are a lot of compounds available, that can be described with information rich high-dimensional representations, but not all of the compounds have information on their toxicity. To address these challenges, we propose semi-supervised learning of feature ranking. The feature rankings are learned in the context of classification and regression as well as in the context of structured output prediction (multi-label classification, hierarchical multi-label classification and multi-target regression). To the best of our knowledge, this is the first work that treats the task of feature ranking within the semi-supervised structured output prediction context. More specifically, we propose two approaches that are based on tree ensembles and the Relief family of algorithms. The extensive evaluation across 38 benchmark datasets reveals the following: Random Forests perform the best for the classification-like tasks, while for the regression-like tasks Extra-PCTs perform the best, Random Forests are the most efficient method considering induction times across all tasks, and semi-supervised feature rankings outperform their supervised counterpart across a majority of the datasets from the different tasks.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/19/2022

Semi-supervised Predictive Clustering Trees for (Hierarchical) Multi-label Classification

Semi-supervised learning (SSL) is a common approach to learning predicti...
research
02/20/2019

Noisy multi-label semi-supervised dimensionality reduction

Noisy labeled data represent a rich source of information that often are...
research
04/06/2021

Semi-supervised empirical Bayes group-regularized factor regression

The features in high dimensional biomedical prediction problems are ofte...
research
11/03/2020

Deep tree-ensembles for multi-output prediction

Recently, deep neural networks have expanded the state-of-art in various...
research
07/27/2020

Oblique Predictive Clustering Trees

Predictive clustering trees (PCTs) are a well established generalization...
research
12/08/2016

Progressive Tree-like Curvilinear Structure Reconstruction with Structured Ranking Learning and Graph Algorithm

We propose a novel tree-like curvilinear structure reconstruction algori...
research
07/06/2018

A Structured Prediction Approach for Label Ranking

We propose to solve a label ranking problem as a structured output regre...

Please sign up or login with your details

Forgot password? Click here to reset