Unique Metric for Health Analysis with Optimization of Clustering Activity and Cross Comparison of Results from Different Approach

10/08/2018
by   Kumarjit Pathak, et al.
8

In machine learning and data mining, Cluster analysis is one of the most widely used unsupervised learning technique. Philosophy of this algorithm is to find similar data items and group them together based on any distance function in multidimensional space. These methods are suitable for finding groups of data that behave in a coherent fashion. The perspective may vary for clustering i.e. the way we want to find similarity, some methods are based on distance such as K-Means technique and some are probability based, like GMM. Understanding prominent segment of data is always challenging as multidimension space does not allow us to have a look and feel of the distance or any visual context on the health of the clustering. While explaining data using clusters, the major problem is to tell how many cluster are good enough to explain the data. Generally basic descriptive statistics are used to estimate cluster behaviour like scree plot, dendrogram etc. We propose a novel method to understand the cluster behaviour which can be used not only to find right number of clusters but can also be used to access the difference of health between different clustering methods on same data. Our technique would also help to also eliminate the noisy variables and optimize the clustering result. keywords - Clustering, Metric, K-means, hierarchical clustering, silhoutte, clustering index, measures

READ FULL TEXT
research
06/05/2018

A Visual Quality Index for Fuzzy C-Means

Cluster analysis is widely used in the areas of machine learning and dat...
research
07/02/2023

Spatiotemporal Cluster Analysis of Gridded Temperature Data – A Comparison Between K-means and MiSTIC

The Earth is a system of numerous interconnected spheres, such as the cl...
research
06/12/2020

Distance-based phylogenetic inference from typing data: a unifying view

Typing methods are widely used in the surveillance of infectious disease...
research
04/27/2023

ClusterNet: A Perception-Based Clustering Model for Scattered Data

Cluster separation in scatterplots is a task that is typically tackled b...
research
08/15/2015

Towards an Axiomatic Approach to Hierarchical Clustering of Measures

We propose some axioms for hierarchical clustering of probability measur...
research
01/23/2017

The Impact of Random Models on Clustering Similarity

Clustering is a central approach for unsupervised learning. After cluste...
research
01/20/2021

Uncovering and Displaying the Coherent Groups of Rank Data by Exploratory Riffle Shuffling

Let n respondents rank order d items, and suppose that d << n. Our main ...

Please sign up or login with your details

Forgot password? Click here to reset