Local intrinsic dimensionality estimators based on concentration of measure

01/31/2020
by   Jonathan Bac, et al.
51

Intrinsic dimensionality (ID) is one of the most fundamental characteristics of multi-dimensional data point clouds. Knowing ID is crucial to choose the appropriate machine learning approach as well as to understand its behavior and validate it. ID can be computed globally for the whole data distribution, or estimated locally in a point. In this paper, we introduce new local estimators of ID based on linear separability of multi-dimensional data point clouds, which is one of the manifestations of concentration of measure. We empirically study the properties of these measures and compare them with other recently introduced ID estimators exploiting various effects of measure concentration. Observed differences in the behaviour of different estimators can be used to anticipate their behaviour in practical applications.

READ FULL TEXT

page 6

page 7

research
01/18/2019

Estimating the effective dimension of large biological datasets using Fisher separability analysis

Modern large-scale datasets are frequently said to be high-dimensional. ...
research
11/05/2021

Boundary Estimation from Point Clouds: Algorithms, Guarantees and Applications

We investigate identifying the boundary of a domain from sample points i...
research
12/06/2018

Observing the Population Dynamics in GE by means of the Intrinsic Dimension

We explore the use of Intrinsic Dimension (ID) for gaining insights in h...
research
09/29/2022

Intrinsic Dimensionality Estimation within Tight Localities: A Theoretical and Experimental Analysis

Accurate estimation of Intrinsic Dimensionality (ID) is of crucial impor...
research
05/16/2022

From Small Scales to Large Scales: Distance-to-Measure Density based Geometric Analysis of Complex Data

How can we tell complex point clouds with different small scale characte...
research
04/05/2023

Local Intrinsic Dimensional Entropy

Most entropy measures depend on the spread of the probability distributi...
research
02/27/2019

Clustering by the local intrinsic dimension: the hidden structure of real-world data

It is well known that a small number of variables is often sufficient to...

Please sign up or login with your details

Forgot password? Click here to reset