Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Fantastic Generalization Measures and Where to Find Them

About

Generalization of deep networks has been of great interest in recent years, resulting in a number of theoretically and empirically motivated complexity measures. However, most papers proposing such measures study only a small set of models, leaving open the question of whether the conclusion drawn from those experiments would remain valid in other settings. We present the first large scale study of generalization in deep networks. We investigate more then 40 complexity measures taken from both theoretical bounds and empirical studies. We train over 10,000 convolutional networks by systematically varying commonly used hyperparameters. Hoping to uncover potentially causal relationships between each measure and generalization, we analyze carefully controlled experiments and show surprising failures of some measures as well as promising measures for further research.

Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, Samy Bengio• 2019

Related benchmarks

TaskDatasetResultRank
Generalization Gap PredictionCIFAR-10
Kendall Rank Correlation0.613
28
Generalization Gap PredictionMNIST
Kendall Correlation0.546
28
Sentiment ClassificationIMDb Shuffled (val)
Kendall's Tau1
6
Biography ClassificationBias-in-Bios
Kendall's tau (Compression)0.33
6
Sentiment ClassificationSentiment Adjectives (val)
Kendall's Tau0.00e+0
6
Generalization Gap Correlation AnalysisCIFAR-10 (test)
Runtime (s)210.3
6
Digit ClassificationColored MNIST (val)
Validation Accuracy (Kendall's tau)0.8
5
Digit ClassificationShuffled MNIST (val)
Accuracy (Kendall's tau)1
5
Generalization Gap PredictionCIFAR-10 and CIFAR-100 grid of 1,152 models (test)
Spearman ρ0.639
5
Showing 9 of 9 rows

Other info

Follow for update