Tradeoffs in Streaming Binary Classification under Limited Inspection Resources

10/05/2021
by   Parisa Hassanzadeh, et al.
5

Institutions are increasingly relying on machine learning models to identify and alert on abnormal events, such as fraud, cyber attacks and system failures. These alerts often need to be manually investigated by specialists. Given the operational cost of manual inspections, the suspicious events are selected by alerting systems with carefully designed thresholds. In this paper, we consider an imbalanced binary classification problem, where events arrive sequentially and only a limited number of suspicious events can be inspected. We model the event arrivals as a non-homogeneous Poisson process, and compare various suspicious event selection methods including those based on static and adaptive thresholds. For each method, we analytically characterize the tradeoff between the minority-class detection rate and the inspection capacity as a function of the data class imbalance and the classifier confidence score densities. We implement the selection methods on a real public fraud detection dataset and compare the empirical results with analytical bounds. Finally, we investigate how class imbalance and the choice of classifier impact the tradeoff.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/05/2021

Statistical Theory for Imbalanced Binary Classification

Within the vast body of statistical theory developed for binary classifi...
research
11/17/2017

Large Neural Network Based Detection of Apnea, Bradycardia and Desaturation Events

Apnea, bradycardia and desaturation (ABD) events often precede life-thre...
research
04/06/2021

Taming Adversarial Robustness via Abstaining

In this work, we consider a binary classification problem and cast it in...
research
05/23/2021

A Study imbalance handling by various data sampling methods in binary classification

The purpose of this research report is to present the our learning curve...
research
01/29/2019

Bayes Imbalance Impact Index: A Measure of Class Imbalanced Dataset for Classification Problem

Recent studies have shown that imbalance ratio is not the only cause of ...
research
07/27/2023

Retrieval-based Text Selection for Addressing Class-Imbalanced Data in Classification

This paper addresses the problem of selecting of a set of texts for anno...
research
10/21/2020

Deep Q-Network-based Adaptive Alert Threshold Selection Policy for Payment Fraud Systems in Retail Banking

Machine learning models have widely been used in fraud detection systems...

Please sign up or login with your details

Forgot password? Click here to reset