Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

One-Shot Coresets: The Case of k-Clustering

About

Scaling clustering algorithms to massive data sets is a challenging task. Recently, several successful approaches based on data summarization methods, such as coresets and sketches, were proposed. While these techniques provide provably good and small summaries, they are inherently problem dependent - the practitioner has to commit to a fixed clustering objective before even exploring the data. However, can one construct small data summaries for a wide range of clustering problems simultaneously? In this work, we affirmatively answer this question by proposing an efficient algorithm that constructs such one-shot summaries for k-clustering problems while retaining strong theoretical guarantees.

Olivier Bachem, Mario Lucic, Silvio Lattanzi• 2017

Related benchmarks

TaskDatasetResultRank
Coreset ConstructionCrime
Wasserstein Distance2.01
30
Coreset Constructiondrug
Wasserstein Distance4.16
30
Coreset ConstructionGerman Credit
Wasserstein Distance0.26
28
Coreset ConstructionAdult
Wasserstein Distance9.37
24
ClassificationCredit Dataset (test)
DD0.07
10
ClassificationDrug Dataset (test)
DD0.16
10
ClassificationCrime Dataset (test)
DD0.45
10
ClassificationAdult Dataset (test)
DD0.12
8
Showing 8 of 8 rows

Other info

Follow for update