Min-Sum Clustering (with Outliers)

11/24/2020
by   Sandip Banerjee, et al.
0

We give a constant factor polynomial time pseudo-approximation algorithm for min-sum clustering with or without outliers. The algorithm is allowed to exclude an arbitrarily small constant fraction of the points. For instance, we show how to compute a solution that clusters 98% of the input data points and pays no more than a constant factor times the optimal solution that clusters 99% of the input data points. More generally, we give the following bicriteria approximation: For any > 0, for any instance with n input points and for any positive integer n'≤ n, we compute in polynomial time a clustering of at least (1-) n' points of cost at most a constant factor greater than the optimal cost of clustering n' points. The approximation guarantee grows with 1/. Our results apply to instances of points in real space endowed with squared Euclidean distance, as well as to points in a metric space, where the number of clusters, and also the dimension if relevant, is arbitrary (part of the input, not an absolute constant).

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/19/2021

Approximation Algorithms For The Euclidean Dispersion Problems

In this article, we consider the Euclidean dispersion problems. Let P={p...
research
03/14/2023

FPT Constant-Approximations for Capacitated Clustering to Minimize the Sum of Cluster Radii

Clustering with capacity constraints is a fundamental problem that attra...
research
05/02/2023

FPT Approximations for Capacitated/Fair Clustering with Outliers

Clustering problems such as k-Median, and k-Means, are motivated from ap...
research
02/11/2023

Partial k-means to avoid outliers, mathematical programming formulations, complexity results

A well-known bottleneck of Min-Sum-of-Square Clustering (MSSC, the celeb...
research
09/02/2023

Approximating Fair k-Min-Sum-Radii in ℝ^d

The k-center problem is a classical clustering problem in which one is a...
research
04/08/2018

Dimensionality's Blessing: Clustering Images by Underlying Distribution

Many high dimensional vector distances tend to a constant. This is typic...
research
01/15/2019

A constant parameterized approximation for hard-capacitated k-means

Hard-capacitated k-means (HCKM) is one of the remaining fundamental prob...

Please sign up or login with your details

Forgot password? Click here to reset