DoGE: Domain Reweighting with Generalization Estimation

About

The coverage and composition of the pretraining data significantly impacts the generalization ability of Large Language Models (LLMs). Despite its importance, recent LLMs still rely on heuristics and trial and error to increase or reduce the influence of data-domains. We propose DOmain reweighting with Generalization Estimation (DoGE), which optimizes the probability of sampling from each domain (domain weights) in a principled way. Our approach is a two-stage process consisting of (i) training a proxy model to obtain domain weights using a bi-level optimization algorithm; (ii) training a larger base model by sampling training domains according to the learned domain weights. In our experiments, we extensively show how DoGE improves the generalization of the base model to any target data mixture. On the SlimPajama dataset, our base model gets better perplexity and few-shot reasoning accuracies across $6$ tasks compared to baseline methods. Moreover, aiming to generalize to out-of-domain target tasks, which is unseen in the pretraining corpus (OOD domain), DoGE can effectively identify inter-domain dependencies, and consistently achieves better test perplexity on the target domain.

Simin Fan, Matteo Pagliardini, Martin Jaggi• 2023

Related benchmarks

Task	Dataset	Result
Commonsense Reasoning	PIQA	Accuracy58.5	757
Commonsense Reasoning	HellaSwag	HellaSwag Accuracy29.2	711
Language Modeling	LAMBADA	Accuracy11.7	412
Common Sense Reasoning	COPA	Accuracy64.5	256
Commonsense Reasoning	OBQA	Accuracy27.1	187
Language Modeling	The Pile	Perplexity2.72	129
Language Modeling	SlimPajama	Perplexity (PPL)3.31	77
Commonsense Reasoning	WinoG	Accuracy50.4	55
Commonsense Reasoning	HellaSwag	Accuracy26.2	47
Reading Comprehension	SciQ	Accuracy60.1	32

Showing 10 of 22 rows

Other info

Follow for update

@wizwand_team Discord