A general framework for implementing distances for categorical variables

01/04/2023
by   Michel van de Velden, et al.
0

The degree to which subjects differ from each other with respect to certain properties measured by a set of variables, plays an important role in many statistical methods. For example, classification, clustering, and data visualization methods all require a quantification of differences in the observed values. We can refer to the quantification of such differences, as distance. An appropriate definition of a distance depends on the nature of the data and the problem at hand. For distances between numerical variables, there exist many definitions that depend on the size of the observed differences. For categorical data, the definition of a distance is more complex, as there is no straightforward quantification of the size of the observed differences. Consequently, many proposals exist that can be used to measure differences based on categorical variables. In this paper, we introduce a general framework that allows for an efficient and transparent implementation of distances between observations on categorical variables. We show that several existing distances can be incorporated into the framework. Moreover, our framework quite naturally leads to the introduction of new distance formulations and allows for the implementation of flexible, case and data specific distance definitions. Furthermore, in a supervised classification setting, the framework can be used to construct distances that incorporate the association between the response and predictor variables and hence improve the performance of distance-based classifiers.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
11/29/2019

Minkowski distances and standardisation for clustering and classification of high dimensional data

There are many distance-based methods for classification and clustering,...
research
12/01/2021

Dimensionality Reduction for Categorical Data

Categorical attributes are those that can take a discrete set of values,...
research
04/08/2021

A global method for mixed categorical optimization with catalogs

In this article, we propose an algorithmic framework for globally solvin...
research
06/27/2019

Evidential distance measure in complex belief function theory

In this paper, an evidential distance measure is proposed which can meas...
research
12/09/2018

A matching based clustering algorithm for categorical data

Cluster analysis is one of the essential tasks in data mining and knowle...
research
09/27/2016

A Transportation L^p Distance for Signal Analysis

Transport based distances, such as the Wasserstein distance and earth mo...
research
01/07/2021

Distances with mixed type variables some modified Gower's coefficients

Nearest neighbor methods have become popular in official statistics, mai...

Please sign up or login with your details

Forgot password? Click here to reset