DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

About

Reinforcement learning (RL) with large language models shows promise in complex reasoning. However, its progress is hindered by the lack of large-scale training data that is sufficiently challenging, contamination-free and verifiable. To this end, we introduce DeepMath-103K, a large-scale mathematical dataset designed with high difficulty (primarily levels 5-9), rigorous decontamination against numerous benchmarks, and verifiable answers for rule-based RL reward. It further includes three distinct R1 solutions adaptable for diverse training paradigms such as supervised fine-tuning (SFT). Spanning a wide range of mathematical topics, DeepMath-103K fosters the development of generalizable and advancing reasoning. Notably, models trained on DeepMath-103K achieve state-of-the-art results on challenging mathematical benchmarks and demonstrate generalization beyond math such as biology, physics and chemistry, underscoring its broad efficacy. Data: https://huggingface.co/datasets/zwhe99/DeepMath-103K.

Zhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu, Xingyu Chen, Yue Wang, Linfeng Song, Dian Yu, Zhenwen Liang, Wenxuan Wang, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, Dong Yu• 2025

Related benchmarks

Task	Dataset	Result
Mathematical Reasoning	AIME 2024	Accuracy34.2	525
Mathematical Reasoning	AIME 2025	Accuracy30	353
Mathematical Reasoning	HMMT 2025	Accuracy11.7	241
Mathematical Reasoning	MATH 500	Accuracy83.4	183
Mathematical Reasoning	Omni-MATH	Accuracy45.4	135
Mathematical Reasoning	Minerva Math	pass@1 Accuracy45.4	104
Mathematical Reasoning	OlympiadBench Math	Accuracy60.2	97
Mathematical Reasoning	AMC 2023	Pass@164.7	67
Mathematical Reasoning	AIME 2025	Accuracy31.7	59
Mathematical Reasoning	Math Reasoning Suite Average	Average Accuracy25.1	49

Showing 10 of 39 rows

Other info

Follow for update

@wizwand_team Discord