Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Sparse Communication for Distributed Gradient Descent

About

We make distributed stochastic gradient descent faster by exchanging sparse updates instead of dense updates. Gradient updates are positively skewed as most updates are near zero, so we map the 99% smallest updates (by absolute value) to zero then exchange sparse matrices. This method can be combined with quantization to further improve the compression. We explore different configurations and apply them to neural machine translation and MNIST image classification tasks. Most configurations work on MNIST, whereas different configurations reduce convergence rate on the more complex translation task. Our experiments show that we can achieve up to 49% speed up on MNIST and 22% on NMT without damaging the final accuracy or BLEU.

Alham Fikri Aji, Kenneth Heafield• 2017

Related benchmarks

TaskDatasetResultRank
Image ClassificationCIFAR100
Accuracy42.2
37
Finance Question AnsweringFinancial-qa-10k
MP55.4
24
General Question Answering & Instruction FollowingDolly
MP Score17.3
24
Medical Question AnsweringMedicalMeadow WikiDoc PatientInfo Med
MP16.8
24
Image ClassificationCIFAR10
Accuracy74.5
7
Showing 5 of 5 rows

Other info

Follow for update