Linear-time Outlier Detection via Sensitivity

05/02/2016
by   Mario Lucic, et al.
0

Outliers are ubiquitous in modern data sets. Distance-based techniques are a popular non-parametric approach to outlier detection as they require no prior assumptions on the data generating distribution and are simple to implement. Scaling these techniques to massive data sets without sacrificing accuracy is a challenging task. We propose a novel algorithm based on the intuition that outliers have a significant influence on the quality of divergence-based clustering solutions. We propose sensitivity - the worst-case impact of a data point on the clustering objective - as a measure of outlierness. We then prove that influence, a (non-trivial) upper-bound on the sensitivity, can be computed by a simple linear time algorithm. To scale beyond a single machine, we propose a communication efficient distributed algorithm. In an extensive experimental evaluation, we demonstrate the effectiveness and establish the statistical significance of the proposed approach. In particular, it outperforms the most popular distance-based approaches while being several orders of magnitude faster.

READ FULL TEXT
research
03/12/2018

Onion-Peeling Outlier Detection in 2-D data Sets

Outlier Detection is a critical and cardinal research task due its array...
research
10/15/2019

MSD-Kmeans: A Novel Algorithm for Efficient Detection of Global and Local Outliers

Outlier detection is a technique in data mining that aims to detect unus...
research
01/05/2018

Clustering with Outlier Removal

Cluster analysis and outlier detection are strongly coupled tasks in dat...
research
10/19/2019

Efficient Discovery of Meaningful Outlier Relationships

We propose PODS (Predictable Outliers in Data-trendS), a method that, gi...
research
03/01/2022

On the impact of outliers in loss reserving

The sensitivity of loss reserving techniques to outliers in the data or ...
research
01/05/2023

Exact and Heuristic Approaches to Speeding Up the MSM Time Series Distance Computation

The computation of the distance of two time series is time-consuming for...
research
02/27/2017

Scalable and Distributed Clustering via Lightweight Coresets

Coresets are compact representations of data sets such that models train...

Please sign up or login with your details

Forgot password? Click here to reset