Robust Clustering with Normal Mixture Models: A Pseudo β-Likelihood Approach

09/10/2020
by   Soumya Chakraborty, et al.
0

As in other estimation scenarios, likelihood based estimation in the normal mixture set-up is highly non-robust against model misspecification and presence of outliers (apart from being an ill-posed optimization problem). We propose a robust alternative to the ordinary likelihood approach for this estimation problem which performs simultaneous estimation and data clustering and leads to subsequent anomaly detection. To invoke robustness, we follow, in spirit, the methodology based on the minimization of the density power divergence (or alternatively, the maximization of the β-likelihood) under suitable constraints. An iteratively reweighted least squares approach has been followed in order to compute our estimators for the component means (or equivalently cluster centers) and component dispersion matrices which leads to simultaneous data clustering. Some exploratory techniques are also suggested for anomaly detection, a problem of great importance in the domain of statistics and machine learning. Existence and consistency of the estimators are established under the aforesaid constraints. We validate our method with simulation studies under different set-ups; it is seen to perform competitively or better compared to the popular existing methods like K-means and TCLUST, especially when the mixture components (i.e., the clusters) share regions with significant overlap or outlying clusters exist with small but non-negligible weights. Two real datasets are also used to illustrate the performance of our method in comparison with others along with an application in image processing. It is observed that our method detects the clusters with lower misclassification rates and successfully points out the outlying (anomalous) observations from these datasets.

READ FULL TEXT

page 31

page 33

research
05/11/2022

Existence and Consistency of the Maximum Pseudo e̱ṯa̱-Likelihood Estimators for Multivariate Normal Mixture Models

Robust estimation under multivariate normal (MVN) mixture model is alway...
research
11/01/2019

Integrated Clustering and Anomaly Detection (INCAD) for Streaming Data (Revised)

Most current clustering based anomaly detection methods use scoring sche...
research
06/19/2019

Robust Clustering Using Tau-Scales

K means is a popular non-parametric clustering procedure introduced by S...
research
02/26/2014

Robust Asymmetric Clustering

Contaminated mixture models are developed for model-based clustering of ...
research
02/13/2021

Robust Model-Based Clustering

We propose a new class of robust and Fisher-consistent estimators for mi...
research
09/16/2020

Clustering Data with Nonignorable Missingness using Semi-Parametric Mixture Models

We are concerned in clustering continuous data sets subject to nonignora...
research
12/21/2021

Anomaly Clustering: Grouping Images into Coherent Clusters of Anomaly Types

We introduce anomaly clustering, whose goal is to group data into semant...

Please sign up or login with your details

Forgot password? Click here to reset