DSGym: A Holistic Framework for Evaluating and Training Data Science Agents

About

Data science agents promise to accelerate discovery and insight-generation by turning data into executable analyses and findings. Yet existing data science benchmarks fall short due to fragmented evaluation interfaces that make cross-benchmark comparison difficult, narrow task coverage and a lack of rigorous data grounding. In particular, we show that a substantial portion of tasks in current benchmarks can be solved without using the actual data. To address these limitations, we introduce DSGym, a standardized framework for evaluating and training data science agents in self-contained execution environments. Unlike static benchmarks, DSGym provides a modular architecture that makes it easy to add tasks, agent scaffolds, and tools, positioning it as a live, extensible testbed. We curate DSGym-Tasks, a holistic task suite that standardizes and refines existing benchmarks via quality and shortcut solvability filtering. We further expand coverage with (1) DSBio: expert-derived bioinformatics tasks grounded in literature and (2) DSPredict: challenging prediction tasks spanning domains such as computer vision, molecular prediction, and single-cell perturbation. Beyond evaluation, DSGym enables agent training via execution-verified data synthesis pipeline. As a case study, we build a 2,000-example training set and trained a 4B model in DSGym that outperforms GPT-4o on standardized analysis benchmarks. Overall, DSGym enables rigorous end-to-end measurement of whether agents can plan, implement, and validate data analyses in realistic scientific context.

Fan Nie, Junlin Wang, Harper Hua, Federico Bianchi, Yongchan Kwon, Zhenting Qi, Owen Queen, Shang Zhu, James Zou• 2026

Related benchmarks

Task	Dataset	Result
Bioinformatics Data Analysis	DSBio	--	13
Data Analysis	QRData Verified	--	13
Data Analysis	DABStep easy	--	13
Data Analysis	DABStep hard	--	13
Data Analysis	DAEval Verified	--	13
Data Prediction	MLEBench lite	--	12
Data Prediction	DSPredict Hard (Private test)	--	12
Data Prediction	DSPredict Easy (test)	--	12

Showing 8 of 8 rows

Other info

Follow for update

@wizwand_team Discord