ATD: Anomalous Topic Discovery in High Dimensional Discrete Data

12/20/2015
by   Hossein Soleimani, et al.
0

We propose an algorithm for detecting patterns exhibited by anomalous clusters in high dimensional discrete data. Unlike most anomaly detection (AD) methods, which detect individual anomalies, our proposed method detects groups (clusters) of anomalies; i.e. sets of points which collectively exhibit abnormal patterns. In many applications this can lead to better understanding of the nature of the atypical behavior and to identifying the sources of the anomalies. Moreover, we consider the case where the atypical patterns exhibit on only a small (salient) subset of the very high dimensional feature space. Individual AD techniques and techniques that detect anomalies using all the features typically fail to detect such anomalies, but our method can detect such instances collectively, discover the shared anomalous patterns exhibited by them, and identify the subsets of salient features. In this paper, we focus on detecting anomalous topics in a batch of text documents, developing our algorithm based on topic models. Results of our experiments show that our method can accurately detect anomalous topics and salient features (words) under each such topic in a synthetic data set and two real-world text corpora and achieves better performance compared to both standard group AD and individual AD techniques. All required code to reproduce our experiments is available from https://github.com/hsoleimani/ATD

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/22/2023

AD-MERCS: Modeling Normality and Abnormality in Unsupervised Anomaly Detection

Most anomaly detection systems try to model normal behavior and assume a...
research
12/05/2022

Prototypical Residual Networks for Anomaly Detection and Localization

Anomaly detection and localization are widely used in industrial manufac...
research
09/30/2014

Data Imputation through the Identification of Local Anomalies

We introduce a comprehensive and statistical framework in a model free s...
research
06/27/2022

Auditing Visualizations: Transparency Methods Struggle to Detect Anomalous Behavior

Transparency methods such as model visualizations provide information th...
research
11/23/2021

Post-discovery Analysis of Anomalous Subsets

Analyzing the behaviour of a population in response to disease and inter...
research
10/18/2018

Unsupervised Anomalous Data Space Specification

Computer algorithms are written with the intent that when run they perfo...
research
10/04/2021

Stochastic functional analysis with applications to robust machine learning

It is well-known that machine learning protocols typically under-utilize...

Please sign up or login with your details

Forgot password? Click here to reset