SSDBCODI: Semi-Supervised Density-Based Clustering with Outliers Detection Integrated

08/10/2022
by   Jiahao Deng, et al.
0

Clustering analysis is one of the critical tasks in machine learning. Traditionally, clustering has been an independent task, separate from outlier detection. Due to the fact that the performance of clustering can be significantly eroded by outliers, a small number of algorithms try to incorporate outlier detection in the process of clustering. However, most of those algorithms are based on unsupervised partition-based algorithms such as k-means. Given the nature of those algorithms, they often fail to deal with clusters of complex, non-convex shapes. To tackle this challenge, we have proposed SSDBCODI, a semi-supervised density-based algorithm. SSDBCODI combines the advantage of density-based algorithms, which are capable of dealing with clusters of complex shapes, with the semi-supervised element, which offers flexibility to adjust the clustering results based on a few user labels. We also merge an outlier detection component with the clustering process. Potential outliers are detected based on three scores generated during the process: (1) reachability-score, which measures how density-reachable a point is to a labeled normal object, (2) local-density-score, which measures the neighboring density of data objects, and (3) similarity-score, which measures the closeness of a point to its nearest labeled outliers. Then in the following step, instance weights are generated for each data instance based on those three scores before being used to train a classifier for further clustering and outlier detection. To enhance the understanding of the proposed algorithm, for our evaluation, we have run our proposed algorithm against some of the state-of-art approaches on multiple datasets and separately listed the results of outlier detection apart from clustering. Our results indicate that our algorithm can achieve superior results with a small percentage of labels.

READ FULL TEXT

page 10

page 11

research
12/01/2019

XGBOD: Improving Supervised Outlier Detection with Unsupervised Representation Learning

A new semi-supervised ensemble algorithm called XGBOD (Extreme Gradient ...
research
01/14/2019

CFOF: A Concentration Free Measure for Anomaly Detection

We present a novel notion of outlier, called the Concentration Free Outl...
research
06/08/2020

Outlier Detection Using a Novel method: Quantum Clustering

We propose a new assumption in outlier detection: Normal data instances ...
research
03/07/2020

RCC-Dual-GAN: An Efficient Approach for Outlier Detection with Few Identified Anomalies

Outlier detection is an important task in data mining and many technolog...
research
12/07/2017

Using SVDD in SimpleMKL for 3D-Shapes Filtering

This paper proposes the adaptation of Support Vector Data Description (S...
research
09/23/2021

Fast Density Estimation for Density-based Clustering Methods

Density-based clustering algorithms are widely used for discovering clus...
research
10/15/2022

D.MCA: Outlier Detection with Explicit Micro-Cluster Assignments

How can we detect outliers, both scattered and clustered, and also expli...

Please sign up or login with your details

Forgot password? Click here to reset