Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Trees from Marginals: Autoregressive drafting with factorized priors

About

Speculative decoding greatly increases the interactivity of autoregressive language models by trading off computation for extra tokens generated in a single forward pass. Factorized draft models are especially efficient because they predict future-token marginals in parallel, but their independence assumption causes acceptance rates to degrade sharply as the speculative budget grows. We analyze this limitation and introduce Weaver, a lightweight autoregressive adapter that constructs proposal trees from the top-K marginals of a factorized drafter. Weaver restores conditional dependencies between proposed tokens while avoiding a full-vocabulary projection. To support fast verification for models with Gated Delta Net layers, we derive a rollback-free tree-verification algorithm and implement optimized CUDA kernels in SGLang. By combining these model and systems contributions we achieve a 4.37-fold speedup over autoregressive decoding, and outperform a highly optimized DFlash baseline by 24.7%.

Yuma Oda, Ryan Mathieu, Roman Knyazhitskiy, Artur Chakhvadze• 2026

Related benchmarks

TaskDatasetResultRank
ChatMT-Bench--
73
MathGSM8K
Speedup4.85
14
MathMATH500
Speedup5.3
14
ChatShareChat
Speedup2.85
8
Aggregate PerformanceMacro Avg.
Speedup4.37
7
CodeHumanEval
Speedup5.17
2
Showing 6 of 6 rows

Other info

Follow for update