Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks

About

Effective training of deep neural networks suffers from two main issues. The first is that the parameter spaces of these models exhibit pathological curvature. Recent methods address this problem by using adaptive preconditioning for Stochastic Gradient Descent (SGD). These methods improve convergence by adapting to the local geometry of parameter space. A second issue is overfitting, which is typically addressed by early stopping. However, recent work has demonstrated that Bayesian model averaging mitigates this problem. The posterior can be sampled by using Stochastic Gradient Langevin Dynamics (SGLD). However, the rapidly changing curvature renders default SGLD methods inefficient. Here, we propose combining adaptive preconditioners with SGLD. In support of this idea, we give theoretical properties on asymptotic convergence and predictive risk. We also provide empirical results for Logistic Regression, Feedforward Neural Nets, and Convolutional Neural Nets, demonstrating that our preconditioned SGLD method gives state-of-the-art performance on these models.

Chunyuan Li, Changyou Chen, David Carlson, Lawrence Carin• 2015

Related benchmarks

TaskDatasetResultRank
Epistemic uncertainty calibrationmaterial plasticity dataset
ECE2.23
14
Image RegressionCells-Tail
TC Score84.92
7
Image RegressionChairAngle Tail
TC Score77.97
7
Image RegressionSkin
TC0.8625
7
Image RegressionCells-Gap
TC0.9114
7
Image RegressionChairAngle-Gap
TC Score92.78
7
Image RegressionCELLS
TC Score94.5
7
Image RegressionChairAngle
TC94.91
7
Image RegressionAerial
TC67.24
7
Age EstimationAPPA-REAL 1.0 (test)
RMSE18.57
3
Showing 10 of 10 rows

Other info

Follow for update