Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
About
We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion's long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Commonsense Reasoning | WinoGrande | Accuracy71.98 | 1581 | |
| Visual Question Answering | ChartQA | Accuracy81.3 | 620 | |
| Multi-discipline Multimodal Understanding | MMMU | Accuracy50.4 | 422 | |
| Visual Question Answering | AI2D | Accuracy75 | 402 | |
| General Knowledge | MMLU | MMLU General Knowledge Accuracy74.68 | 373 | |
| Visual Question Answering | RealworldQA | Accuracy62.6 | 327 | |
| Mathematical Reasoning | Minerva Math | Accuracy67.38 | 251 | |
| Commonsense Reasoning | ARC-E | Accuracy83.38 | 249 | |
| Math Reasoning | GSM8K | Accuracy (GSM8K)88.48 | 190 | |
| Commonsense Reasoning | HellaSwag | Accuracy76.08 | 106 |