Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CodeAlchemy: Synthetic Code Rewriting at Scale

About

Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats. While synthetic data has proven transformative for language models, code remains largely unexplored beyond limited quality improvements. We present CodeAlchemy, a synthetic data generation framework that transforms publicly sourced code into semantically-rich training data through 5 strategies: CodeEnhance (quality-aware rewriting), CodeQA (template-based problems), CodeDev (developer tasks), CodeDialogue (multi-turn conversations), and CodeTrace (execution traces). We process 3 corpora across 15 languages to generate 500B+ tokens of synthetic data plus 350B reasoning tokens, orders of magnitude more than prior efforts. CodeTrace instruments and executes 1.3M+ files across 14 languages and 5K libraries, capturing control flow, state tracking, and library knowledge. We introduce DevEval (developer tasks) and TraceEval (execution prediction) benchmarks; frontier models like Claude Sonnet 4.5 achieve only 5.6% exact match on TraceEval, revealing critical gaps in semantic understanding. Our 3B models achieve 83.5% on HumanEval, 63.2% on MBPP, 8.09% win rate on DevEval, and 15.36 ROUGE-2 on TraceEval, outperforming frontier models 10x the size including 27B Gemma-3 and 32B Granite-4.0.

Ankit Gupta, Aditya Prasad, Rameswar Panda• 2026

Related benchmarks

TaskDatasetResultRank
Code GenerationHumanEval+
Average Score (AVG)43.3
27
Code GenerationMBPP+
MB+ Score54.2
23
Code GenerationHumanEval
HumanEval Score48.7
17
Code GenerationMBPP
MB Score65
17
Code Reasoning (Input Prediction)CRUXEval-I
CX-I35.7
10
Code Reasoning (Output Prediction)CRUXEval-O
CX-O Score36.4
10
Developer Knowledge EvaluationDevEval--
7
Code trace predictionTraceEval stack-edu (held-out)--
5
Showing 8 of 8 rows

Other info

Follow for update