A geometric framework for outlier detection in high-dimensional data

07/01/2022
by   Moritz Herrmann, et al.
0

Outlier or anomaly detection is an important task in data analysis. We discuss the problem from a geometrical perspective and provide a framework that exploits the metric structure of a data set. Our approach rests on the manifold assumption, i.e., that the observed, nominally high-dimensional data lie on a much lower dimensional manifold and that this intrinsic structure can be inferred with manifold learning methods. We show that exploiting this structure significantly improves the detection of outlying observations in high-dimensional data. We also suggest a novel, mathematically precise, and widely applicable distinction between distributional and structural outliers based on the geometry and topology of the data manifold that clarifies conceptual ambiguities prevalent throughout the literature. Our experiments focus on functional data as one class of structured high-dimensional data, but the framework we propose is completely general and we include image and graph data applications. Our results show that the outlier structure of high-dimensional and non-tabular data can be detected and visualized using manifold learning methods and quantified using standard outlier scoring methods applied to the manifold embedding vectors.

READ FULL TEXT

page 5

page 10

research
09/14/2021

A geometric perspective on functional outlier detection

We consider functional outlier detection from a geometric perspective, s...
research
09/09/2019

Outlier Detection in High Dimensional Data

High-dimensional data poses unique challenges in outlier detection proce...
research
12/28/2016

Optimal bandwidth estimation for a fast manifold learning algorithm to detect circular structure in high-dimensional data

We provide a way to infer about existence of topological circularity in ...
research
08/16/2020

Geometric Foundations of Data Reduction

The purpose of this paper is to write a complete survey of the (spectral...
research
03/03/2021

Detecting Outliers in High-dimensional Data with Mixed Variable Types using Conditional Gaussian Regression Models

Outlier detection has gained increasing interest in recent years, due to...
research
04/03/2021

Joint Geometric and Topological Analysis of Hierarchical Datasets

In a world abundant with diverse data arising from complex acquisition t...
research
02/14/2008

FINE: Fisher Information Non-parametric Embedding

We consider the problems of clustering, classification, and visualizatio...

Please sign up or login with your details

Forgot password? Click here to reset