Linear-time Outlier Detection via Sensitivity

05/02/2016
by   Mario Lucic, et al.
0

Outliers are ubiquitous in modern data sets. Distance-based techniques are a popular non-parametric approach to outlier detection as they require no prior assumptions on the data generating distribution and are simple to implement. Scaling these techniques to massive data sets without sacrificing accuracy is a challenging task. We propose a novel algorithm based on the intuition that outliers have a significant influence on the quality of divergence-based clustering solutions. We propose sensitivity - the worst-case impact of a data point on the clustering objective - as a measure of outlierness. We then prove that influence, a (non-trivial) upper-bound on the sensitivity, can be computed by a simple linear time algorithm. To scale beyond a single machine, we propose a communication efficient distributed algorithm. In an extensive experimental evaluation, we demonstrate the effectiveness and establish the statistical significance of the proposed approach. In particular, it outperforms the most popular distance-based approaches while being several orders of magnitude faster.

READ FULL TEXT
03/12/2018

Onion-Peeling Outlier Detection in 2-D data Sets

Outlier Detection is a critical and cardinal research task due its array...
10/15/2019

MSD-Kmeans: A Novel Algorithm for Efficient Detection of Global and Local Outliers

Outlier detection is a technique in data mining that aims to detect unus...
01/05/2018

Clustering with Outlier Removal

Cluster analysis and outlier detection are strongly coupled tasks in dat...
10/19/2019

Efficient Discovery of Meaningful Outlier Relationships

We propose PODS (Predictable Outliers in Data-trendS), a method that, gi...
03/01/2022

On the impact of outliers in loss reserving

The sensitivity of loss reserving techniques to outliers in the data or ...
02/27/2017

Scalable and Distributed Clustering via Lightweight Coresets

Coresets are compact representations of data sets such that models train...
10/06/2019

Fast Detection of Outliers in Data Streams with the Q_n Estimator

We present FQN (Fast Q_n), a novel algorithm for fast detection of outli...