KnAC: an approach for enhancing cluster analysis with background knowledge and explanations

12/16/2021
by   Szymon Bobek, et al.
0

Pattern discovery in multidimensional data sets has been a subject of research since decades. There exists a wide spectrum of clustering algorithms that can be used for that purpose. However, their practical applications share in common the post-clustering phase, which concerns expert-based interpretation and analysis of the obtained results. We argue that this can be a bottleneck of the process, especially in the cases where domain knowledge exists prior to clustering. Such a situation requires not only a proper analysis of automatically discovered clusters, but also a conformance checking with existing knowledge. In this work, we present Knowledge Augmented Clustering (KnAC), which main goal is to confront expert-based labelling with automated clustering for the sake of updating and refining the former. Our solution does not depend on any ready clustering algorithm, nor introduce one. Instead KnAC can serve as an augmentation of an arbitrary clustering algorithm, making the approach robust and model-agnostic. We demonstrate the feasibility of our method on artificially, reproducible examples and on a real life use case scenario.

READ FULL TEXT

page 9

page 21

page 26

page 27

research
12/09/2018

A matching based clustering algorithm for categorical data

Cluster analysis is one of the essential tasks in data mining and knowle...
research
02/07/2021

A self-adaptive and robust fission clustering algorithm via heat diffusion and maximal turning angle

Cluster analysis, which focuses on the grouping and categorization of si...
research
06/11/2020

Automating Cluster Analysis to Generate Customer Archetypes for Residential Energy Consumers in South Africa

Time series clustering is frequently used in the energy domain to genera...
research
08/21/2020

ConiVAT: Cluster Tendency Assessment and Clustering with Partial Background Knowledge

The VAT method is a visual technique for determining the potential clust...
research
10/22/2019

Multiple Sample Clustering

The clustering algorithms that view each object data as a single sample ...
research
10/14/2022

GriT-DBSCAN: A Spatial Clustering Algorithm for Very Large Databases

DBSCAN is a fundamental spatial clustering algorithm with numerous pract...
research
03/29/2021

Automatic Clustering in Hyrise

Physical data layout is an important performance factor for modern datab...

Please sign up or login with your details

Forgot password? Click here to reset