Interval-based Prediction Uncertainty Bound Computation in Learning with Missing Values

03/01/2018
by   Hiroyuki Hanada, et al.
0

The problem of machine learning with missing values is common in many areas. A simple approach is to first construct a dataset without missing values simply by discarding instances with missing entries or by imputing a fixed value for each missing entry, and then train a prediction model with the new dataset. A drawback of this naive approach is that the uncertainty in the missing entries is not properly incorporated in the prediction. In order to evaluate prediction uncertainty, the multiple imputation (MI) approach has been studied, but the performance of MI is sensitive to the choice of the probabilistic model of the true values in the missing entries, and the computational cost of MI is high because multiple models must be trained. In this paper, we propose an alternative approach called the Interval-based Prediction Uncertainty Bounding (IPUB) method. The IPUB method represents the uncertainties due to missing entries as intervals, and efficiently computes the lower and upper bounds of the prediction results when all possible training sets constructed by imputing arbitrary values in the intervals are considered. The IPUB method can be applied to a wide class of convex learning algorithms including penalized least-squares regression, support vector machine (SVM), and logistic regression. We demonstrate the advantages of the IPUB method by comparing it with an existing method in numerical experiment with benchmark datasets.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
04/07/2016

Multilevel Weighted Support Vector Machine for Classification on Healthcare Data with Missing Values

This work is motivated by the needs of predictive analytics on healthcar...
research
03/05/2019

What to Expect of Classifiers? Reasoning about Logistic Regression with Missing Features

While discriminative classifiers often yield strong predictive performan...
research
10/26/2022

Imputation of missing values in multi-view data

When missing values occur in multi-view data, all features in a view are...
research
03/15/2017

Aggregation of Classifiers: A Justifiable Information Granularity Approach

In this study, we introduce a new approach to combine multi-classifiers ...
research
06/18/2020

Matrix Completion with Quantified Uncertainty through Low Rank Gaussian Copula

Modern large scale datasets are often plagued with missing entries; inde...
research
08/15/2019

Combining Prediction Intervals on Multi-Source Non-Disclosed Regression Datasets

Conformal Prediction is a framework that produces prediction intervals b...
research
09/06/2017

Clustering of Data with Missing Entries using Non-convex Fusion Penalties

The presence of missing entries in data often creates challenges for pat...

Please sign up or login with your details

Forgot password? Click here to reset