intRinsic: an R package for model-based estimation of the intrinsic dimension of a dataset

02/23/2021
by   Francesco Denti, et al.
0

The estimation of the intrinsic dimension of a dataset is a fundamental step in most dimensionality reduction techniques. This article illustrates intRinsic, an R package that implements novel state-of-the-art likelihood-based estimators of the intrinsic dimension of a dataset. In detail, the methods included in this package are the TWO-NN, Gride, and Hidalgo models. To allow these novel estimators to be easily accessible, the package contains a few high-level, intuitive functions that rely on a broader set of efficient, low-level routines. intRinsic encompasses models that fall into two categories: homogeneous and heterogeneous intrinsic dimension estimators. The first category contains the TWO-NN and Gride models. The functions dedicated to these two methods carry out inference under both the frequentist and Bayesian frameworks. In the second category we find Hidalgo, a Bayesian mixture model, for which an efficient Gibbs sampler is implemented. After discussing the theoretical background, we demonstrate the performance of the models on simulated datasets. This way, we can assess the results by comparing them with the ground truth. Then, we employ the package to study the intrinsic dimension of the Alon dataset, obtained from a famous microarray experiment. We show how the estimation of homogeneous and heterogeneous intrinsic dimensions allows us to gain valuable insights about the topological structure of a dataset.

READ FULL TEXT

page 27

page 31

page 32

page 34

page 35

research
04/28/2021

Distributional Results for Model-Based Intrinsic Dimension Estimators

Modern datasets are characterized by a large number of features that may...
research
05/22/2020

Rdimtools: An R package for Dimension Reduction and Intrinsic Dimension Estimation

Discovering patterns of the complex high-dimensional data is a long-stan...
research
05/30/2020

direpack: A Python 3 package for state-of-the-art statistical dimension reduction methods

The direpack package aims to establish a set of modern statistical dimen...
research
03/15/2012

Regularized Maximum Likelihood for Intrinsic Dimension Estimation

We propose a new method for estimating the intrinsic dimension of a data...
research
10/19/2011

k-NN Regression Adapts to Local Intrinsic Dimension

Many nonparametric regressors were recently shown to converge at rates t...
research
02/11/2020

The role of intrinsic dimension in high-resolution player tracking data – Insights in basketball

A new range of statistical analysis has emerged in sports after the intr...
research
09/29/2022

Intrinsic Dimensionality Estimation within Tight Localities: A Theoretical and Experimental Analysis

Accurate estimation of Intrinsic Dimensionality (ID) is of crucial impor...

Please sign up or login with your details

Forgot password? Click here to reset