Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

An Improved Analysis of Stochastic Gradient Descent with Momentum

About

SGD with momentum (SGDM) has been widely applied in many machine learning tasks, and it is often applied with dynamic stepsizes and momentum weights tuned in a stagewise manner. Despite of its empirical advantage over SGD, the role of momentum is still unclear in general since previous analyses on SGDM either provide worse convergence bounds than those of SGD, or assume Lipschitz or quadratic objectives, which fail to hold in practice. Furthermore, the role of dynamic parameters has not been addressed. In this work, we show that SGDM converges as fast as SGD for smooth objectives under both strongly convex and nonconvex settings. We also establish \textit{the first} convergence guarantee for the multistage setting, and show that the multistage strategy is beneficial for SGDM compared to using fixed parameters. Finally, we verify these theoretical claims by numerical experiments.

Yanli Liu, Yuan Gao, Wotao Yin• 2020

Related benchmarks

TaskDatasetResultRank
Remaining Useful Life predictionC-MAPSS FD002
RMSE43.47
88
Remaining Useful Life predictionC-MAPSS FD003
RMSE32.6
84
Remaining Useful Life predictionC-MAPSS FD004
RMSE51.23
76
Remaining Useful Life predictionC-MAPSS FD001 (test)
RMSE31.05
32
Remaining Useful Life predictionPredictive Maintenance (test)
MSE885.3
8
Showing 5 of 5 rows

Other info

Follow for update