Validation of cluster analysis results on validation data: A systematic framework

03/01/2021
by   Theresa Ullmann, et al.
0

Cluster analysis refers to a wide range of data analytic techniques for class discovery and is popular in many application fields. To judge the quality of a clustering result, different cluster validation procedures have been proposed in the literature. While there is extensive work on classical validation techniques, such as internal and external validation, less attention has been given to validating and replicating a clustering result using a validation dataset. Such a dataset may be part of the original dataset, which is separated before analysis begins, or it could be an independently collected dataset. We present a systematic structured framework for validating clustering results on validation data that includes most existing validation approaches. In particular, we review classical validation techniques such as internal and external validation, stability analysis, hypothesis testing, and visual validation, and show how they can be interpreted in terms of our framework. We precisely define and formalise different types of validation of clustering results on a validation dataset and explain how each type can be implemented in practice. Furthermore, we give examples of how clustering studies from the applied literature that used a validation dataset can be classified into the framework.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/20/2022

Sanity Check for External Clustering Validation Benchmarks using Internal Validation Measures

We address the lack of reliability in benchmarking clustering techniques...
research
08/03/2021

Impact of Load Demand Dataset Characteristics on Clustering Validation Indices

With the inclusion of smart meters, electricity load consumption data ca...
research
08/27/2020

reval: a Python package to determine the best number of clusters with stability-based relative clustering validation

Determining the number of clusters that best partitions a dataset can be...
research
06/24/2021

A review of systematic selection of clustering algorithms and their evaluation

Data analysis plays an indispensable role for value creation in industry...
research
04/04/2023

Clustering Validation with The Area Under Precision-Recall Curves

Confusion matrices and derived metrics provide a comprehensive framework...
research
06/24/2020

Spatial Pattern Recognition with Adjacency-Clustering: Improved Diagnostics for Semiconductor Wafer Bin Maps

In semiconductor manufacturing, statistical quality control hinges on an...
research
11/26/2017

Visual Subpopulation Discovery and Validation in Cohort Study Data

Epidemiology aims at identifying subpopulations of cohort participants t...

Please sign up or login with your details

Forgot password? Click here to reset