Model-based clustering of partial records

03/30/2021
by   Emily M. Goren, et al.
0

Partially recorded data are frequently encountered in many applications and usually clustered by first removing incomplete cases or features with missing values, or by imputing missing values, followed by application of a clustering algorithm to the resulting altered dataset. Here, we develop clustering methodology through a model-based approach using the marginal density for the observed values, assuming a finite mixture model of multivariate t distributions. We compare our approximate algorithm to the corresponding full expectation-maximization (EM) approach that considers the missing values in the incomplete data set and makes a missing at random (MAR) assumption, as well as case deletion and imputation methods. Since only the observed values are utilized, our approach is computationally more efficient than imputation or full EM. Simulation studies demonstrate that our approach has favorable recovery of the true cluster partition compared to case deletion and imputation under various missingness mechanisms, and is at least competitive with the full EM approach, even when MAR assumptions are violated. Our methodology is demonstrated on a problem of clustering gamma-ray bursts and is implemented at https://github.com/emilygoren/MixtClust.

READ FULL TEXT

page 6

page 8

page 9

research
12/20/2021

Model-based Clustering with Missing Not At Random Data

In recent decades, technological advances have made it possible to colle...
research
06/08/2021

Clustering with missing data: which imputation model for which cluster analysis method?

Multiple imputation (MI) is a popular method for dealing with missing va...
research
02/26/2019

Optimal Clustering with Missing Values

Missing values frequently arise in modern biomedical studies due to vari...
research
04/14/2018

Simultaneous Edit and Imputation for Household Data with Structural Zeros

Multivariate categorical data nested within households often include rep...
research
12/29/2018

Imputation and low-rank estimation with Missing Non At Random data

Missing values challenge data analysis because many supervised and unsu-...
research
05/25/2017

Fast Causal Inference with Non-Random Missingness by Test-Wise Deletion

Many real datasets contain values missing not at random (MNAR). In this ...
research
06/20/2020

Scalable Identification of Partially Observed Systems with Certainty-Equivalent EM

System identification is a key step for model-based control, estimator d...

Please sign up or login with your details

Forgot password? Click here to reset