Scalable Clustering: Large Scale Unsupervised Learning of Gaussian Mixture Models with Outliers

02/28/2023
by   Yijia Zhou, et al.
0

Clustering is a widely used technique with a long and rich history in a variety of areas. However, most existing algorithms do not scale well to large datasets, or are missing theoretical guarantees of convergence. This paper introduces a provably robust clustering algorithm based on loss minimization that performs well on Gaussian mixture models with outliers. It provides theoretical guarantees that the algorithm obtains high accuracy with high probability under certain assumptions. Moreover, it can also be used as an initialization strategy for k-means clustering. Experiments on real-world large-scale datasets demonstrate the effectiveness of the algorithm when clustering a large number of clusters, and a k-means algorithm initialized by the algorithm outperforms many of the classic clustering methods in both speed and accuracy, while scaling well to large datasets such as ImageNet.

READ FULL TEXT

page 14

page 15

page 17

research
04/08/2018

Unsupervised Learning of Mixture Models with a Uniform Background Component

Gaussian Mixture Models are one of the most studied and mature models in...
research
06/16/2023

Adversarially robust clustering with optimality guarantees

We consider the problem of clustering data points coming from sub-Gaussi...
research
06/10/2015

Fast Online Clustering with Randomized Skeleton Sets

We present a new fast online clustering algorithm that reliably recovers...
research
05/02/2016

Tradeoffs for Space, Time, Data and Risk in Unsupervised Learning

Faced with massive data, is it possible to trade off (statistical) risk,...
research
09/01/2023

Consistency of Lloyd's Algorithm Under Perturbations

In the context of unsupervised learning, Lloyd's algorithm is one of the...
research
12/27/2020

Generalized Categorisation of Digital Pathology Whole Image Slides using Unsupervised Learning

This project aims to break down large pathology images into small tiles ...
research
09/28/2016

StruClus: Structural Clustering of Large-Scale Graph Databases

We present a structural clustering algorithm for large-scale datasets of...

Please sign up or login with your details

Forgot password? Click here to reset