Towards Reliable Rare Category Analysis on Graphs via Individual Calibration

07/19/2023
by   Longfeng Wu, et al.
0

Rare categories abound in a number of real-world networks and play a pivotal role in a variety of high-stakes applications, including financial fraud detection, network intrusion detection, and rare disease diagnosis. Rare category analysis (RCA) refers to the task of detecting, characterizing, and comprehending the behaviors of minority classes in a highly-imbalanced data distribution. While the vast majority of existing work on RCA has focused on improving the prediction performance, a few fundamental research questions heretofore have received little attention and are less explored: How confident or uncertain is a prediction model in rare category analysis? How can we quantify the uncertainty in the learning process and enable reliable rare category analysis? To answer these questions, we start by investigating miscalibration in existing RCA methods. Empirical results reveal that state-of-the-art RCA methods are mainly over-confident in predicting minority classes and under-confident in predicting majority classes. Motivated by the observation, we propose a novel individual calibration framework, named CALIRARE, for alleviating the unique challenges of RCA, thus enabling reliable rare category analysis. In particular, to quantify the uncertainties in RCA, we develop a node-level uncertainty quantification algorithm to model the overlapping support regions with high uncertainty; to handle the rarity of minority classes in miscalibration calculation, we generalize the distribution-based calibration metric to the instance level and propose the first individual calibration measurement on graphs named Expected Individual Calibration Error (EICE). We perform extensive experimental evaluations on real-world datasets, including rare category characterization and model calibration tasks, which demonstrate the significance of our proposed framework.

READ FULL TEXT
research
06/06/2023

Effective Intrusion Detection in Highly Imbalanced IoT Networks with Lightweight S2CGAN-IDS

Since the advent of the Internet of Things (IoT), exchanging vast amount...
research
05/20/2023

Distribution-Free Model-Agnostic Regression Calibration via Nonparametric Methods

In this paper, we consider the uncertainty quantification problem for re...
research
08/14/2020

Rb-PaStaNet: A Few-Shot Human-Object Interaction Detection Based on Rules and Part States

Existing Human-Object Interaction (HOI) Detection approaches have achiev...
research
08/16/2023

Dual-Branch Temperature Scaling Calibration for Long-Tailed Recognition

The calibration for deep neural networks is currently receiving widespre...
research
08/16/2021

TL-SDD: A Transfer Learning-Based Method for Surface Defect Detection with Few Samples

Surface defect detection plays an increasingly important role in manufac...
research
01/29/2019

Rare geometries: revealing rare categories via dimension-driven statistics

In many situations, the classes of data points of primary interest also ...
research
06/28/2019

Continual Rare-Class Recognition with Emerging Novel Subclasses

Given a labeled dataset that contains a rare (or minority) class of of-i...

Please sign up or login with your details

Forgot password? Click here to reset