Multiple Scaled Contaminated Normal Distribution and Its Application in Clustering

10/21/2018
by   Antonio Punzo, et al.
0

The multivariate contaminated normal (MCN) distribution represents a simple heavy-tailed generalization of the multivariate normal (MN) distribution to model elliptical contoured scatters in the presence of mild outliers, referred to as "bad" points. The MCN can also automatically detect bad points. The price of these advantages is two additional parameters, both with specific and useful interpretations: proportion of good observations and degree of contamination. However, points may be bad in some dimensions but good in others. The use of an overall proportion of good observations and of an overall degree of contamination is limiting. To overcome this limitation, we propose a multiple scaled contaminated normal (MSCN) distribution with a proportion of good observations and a degree of contamination for each dimension. Once the model is fitted, each observation has a posterior probability of being good with respect to each dimension. Thanks to this probability, we have a method for simultaneous directional robust estimation of the parameters of the MN distribution based on down-weighting and for the automatic directional detection of bad points by means of maximum a posteriori probabilities. The term "directional" is added to specify that the method works separately for each dimension. Mixtures of MSCN distributions are also proposed as an application of the proposed model for robust clustering. An extension of the EM algorithm is used for parameter estimation based on the maximum likelihood approach. Real and simulated data are used to show the usefulness of our mixture with respect to well-established mixtures of symmetric distributions with heavy tails.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/08/2020

Mixtures of Contaminated Matrix Variate Normal Distributions

Analysis of three-way data is becoming ever more prevalent in the litera...
research
02/26/2014

Robust Asymmetric Clustering

Contaminated mixture models are developed for model-based clustering of ...
research
06/22/2020

The multivariate tail-inflated normal distribution and its application in finance

This paper introduces the multivariate tail-inflated normal (MTIN) distr...
research
07/14/2017

A new look at the inverse Gaussian distribution

The inverse Gaussian (IG) is one of the most famous and considered distr...
research
10/16/2020

Robust Estimation for Multivariate Wrapped Models

A weighted likelihood technique for robust estimation of a multivariate ...
research
12/10/2021

On the identification of the riskiest directional components from multivariate heavy-tailed data

In univariate data, there exist standard procedures for identifying domi...
research
06/03/2019

Unconstrained representation of orthogonal matrices with application to common principle components

Many statistical problems involve the estimation of a (d× d) orthogonal ...

Please sign up or login with your details

Forgot password? Click here to reset