Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

DSGym: A Holistic Framework for Evaluating and Training Data Science Agents

About

Data science agents promise to accelerate discovery and insight-generation by turning data into executable analyses and findings. Yet existing data science benchmarks fall short due to fragmented evaluation interfaces that make cross-benchmark comparison difficult, narrow task coverage and a lack of rigorous data grounding. In particular, we show that a substantial portion of tasks in current benchmarks can be solved without using the actual data. To address these limitations, we introduce DSGym, a standardized framework for evaluating and training data science agents in self-contained execution environments. Unlike static benchmarks, DSGym provides a modular architecture that makes it easy to add tasks, agent scaffolds, and tools, positioning it as a live, extensible testbed. We curate DSGym-Tasks, a holistic task suite that standardizes and refines existing benchmarks via quality and shortcut solvability filtering. We further expand coverage with (1) DSBio: expert-derived bioinformatics tasks grounded in literature and (2) DSPredict: challenging prediction tasks spanning domains such as computer vision, molecular prediction, and single-cell perturbation. Beyond evaluation, DSGym enables agent training via execution-verified data synthesis pipeline. As a case study, we build a 2,000-example training set and trained a 4B model in DSGym that outperforms GPT-4o on standardized analysis benchmarks. Overall, DSGym enables rigorous end-to-end measurement of whether agents can plan, implement, and validate data analyses in realistic scientific context.

Fan Nie, Junlin Wang, Harper Hua, Federico Bianchi, Yongchan Kwon, Zhenting Qi, Owen Queen, Shang Zhu, James Zou• 2026

Related benchmarks

TaskDatasetResultRank
Bioinformatics Data AnalysisDSBio--
13
Data AnalysisQRData Verified--
13
Data AnalysisDABStep easy--
13
Data AnalysisDABStep hard--
13
Data AnalysisDAEval Verified--
13
Data PredictionMLEBench lite--
12
Data PredictionDSPredict Hard (Private test)--
12
Data PredictionDSPredict Easy (test)--
12
Showing 8 of 8 rows

Other info

Follow for update