Deep Embedded K-Means Clustering

About

Recently, deep clustering methods have gained momentum because of the high representational power of deep neural networks (DNNs) such as autoencoder. The key idea is that representation learning and clustering can reinforce each other: Good representations lead to good clustering while good clustering provides good supervisory signals to representation learning. Critical questions include: 1) How to optimize representation learning and clustering? 2) Should the reconstruction loss of autoencoder be considered always? In this paper, we propose DEKM (for Deep Embedded K-Means) to answer these two questions. Since the embedding space generated by autoencoder may have no obvious cluster structures, we propose to further transform the embedding space to a new space that reveals the cluster-structure information. This is achieved by an orthonormal transformation matrix, which contains the eigenvectors of the within-class scatter matrix of K-means. The eigenvalues indicate the importance of the eigenvectors' contributions to the cluster-structure information in the new space. Our goal is to increase the cluster-structure information. To this end, we discard the decoder and propose a greedy method to optimize the representation. Representation learning and clustering are alternately optimized by DEKM. Experimental results on the real-world datasets demonstrate that DEKM achieves state-of-the-art performance.

Wengang Guo, Kaiyan Lin, Wei Ye• 2021

Related benchmarks

Task	Dataset	Result
Clustering	MNIST	NMI0.9106	113
Clustering	USPS	NMI82.23	104
Clustering	COIL-20	ACC72.62	47
Clustering	REUTERS 10K	ACC76.42	37
Deep Clustering	Fashion MNIST (test)	SC0.819	28
Clustering	FRGC	NMI0.5078	22
Deep Clustering	CIFAR-100	SC0.047	21
Deep Clustering	STL-10	SC0.804	19
Deep Clustering	CIFAR-10	SC0.622	18
Deep Clustering	USPS	SC0.843	14

Showing 10 of 15 rows

Other info

Code

Follow for update

@wizwand_team Discord