bigMap: Big Data Mapping with Parallelized t-SNE

12/24/2018
by   Joan Garriga, et al.
0

We introduce an improved unsupervised clustering protocol specially suited for large-scale structured data. The protocol follows three steps: a dimensionality reduction of the data, a density estimation over the low dimensional representation of the data, and a final segmentation of the density landscape. For the dimensionality reduction step we introduce a parallelized implementation of the well-known t-Stochastic Neighbouring Embedding (t-SNE) algorithm that significantly alleviates some inherent limitations, while improving its suitability for large datasets. We also introduce a new adaptive Kernel Density Estimation particularly coupled with the t-SNE framework in order to get accurate density estimates out of the embedded data, and a variant of the rainfalling watershed algorithm to identify clusters within the density landscape. The whole mapping protocol is wrapped in the bigMap R package, together with visualization and analysis tools to ease the qualitative and quantitative assessment of the clustering.

READ FULL TEXT

page 7

page 15

page 20

page 22

research
01/31/2022

Topology-Preserving Dimensionality Reduction via Interleaving Optimization

Dimensionality reduction techniques are powerful tools for data preproce...
research
11/07/2018

SRP: Efficient class-aware embedding learning for large-scale data via supervised random projections

Supervised dimensionality reduction strategies have been of great intere...
research
05/06/2018

Branching embedding: A heuristic dimensionality reduction algorithm based on hierarchical clustering

This paper proposes a new dimensionality reduction algorithm named branc...
research
04/28/2014

Conditional Density Estimation with Dimensionality Reduction via Squared-Loss Conditional Entropy Minimization

Regression aims at estimating the conditional mean of output given input...
research
09/29/2015

Foundations of Coupled Nonlinear Dimensionality Reduction

In this paper we introduce and analyze the learning scenario of coupled ...
research
08/15/2019

Analyzing the Fine Structure of Distributions

One aim of data mining is the identification of interesting structures i...
research
09/13/2017

Visualization of Big Spatial Data using Coresets for Kernel Density Estimates

The size of large, geo-located datasets has reached scales where visuali...

Please sign up or login with your details

Forgot password? Click here to reset