Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training

About

Pre-training datasets are typically collected from web content and lack inherent domain divisions. For instance, widely used datasets like Common Crawl do not include explicit domain labels, while manually curating labeled datasets such as The Pile is labor-intensive. Consequently, identifying an optimal pre-training data mixture remains a challenging problem, despite its significant benefits for pre-training performance. To address these challenges, we propose CLustering-based Iterative Data Mixture Bootstrapping (Nemotron-CLIMB), an automated framework that discovers, evaluates, and refines data mixtures in a pre-training setting. Specifically, Nemotron-CLIMB embeds and clusters large-scale datasets in a semantic space and then iteratively searches for optimal mixtures using a smaller proxy model and a predictor. When continuously trained on 400B tokens with this mixture, our 1B model exceeds the state-of-the-art Llama-3.2-1B by 2.0%. Moreover, we observe that optimizing for a specific domain (e.g., Social Sciences) yields a 5% improvement over random sampling. Finally, we introduce Nemotron-ClimbLab, a filtered 1.2-trillion-token corpus with 20 clusters as a research playground, and Nemotron-ClimbMix, a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. We analyze the final data mixture, elucidating the characteristics of an optimal data mixture. Our data is available at: https://research.nvidia.com/labs/lpr/climb/

Shizhe Diao, Yu Yang, Yonggan Fu, Xin Dong, Dan Su, Markus Kliegl, Zijia Chen, Peter Belcak, Yoshi Suhara, Hongxu Yin, Mostofa Patwary, Yingyan Lin, Jan Kautz, Pavlo Molchanov• 2025

Related benchmarks

TaskDatasetResultRank
Commonsense ReasoningHellaSwag
HellaSwag Accuracy66.9
897
Reading ComprehensionBoolQ
Accuracy (BoolQ)66.02
258
Word PredictionLAMBADA
Accuracy59.05
222
Social Commonsense ReasoningSocialIQA
Accuracy43.55
150
Science Question AnsweringARC Easy
Accuracy73.57
108
Commonsense ReasoningWinoGrande
Accuracy63.54
94
Science Question AnsweringOpenBookQA
Accuracy41.2
89
Multi-task Language UnderstandingMMLU
Top-1 Accuracy36.47
46
Truthful Question AnsweringTruthfulQA
Accuracy (TruthfulQA)39.06
37
Reading ComprehensionRACE
Accuracy (RACE)36.65
26
Showing 10 of 10 rows

Other info

Follow for update