Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Adaptive Block Diffusion: Resolving Training-Inference Mismatch in Diffusion Language Models

About

Diffusion Language Models (DLMs) are typically trained under fixed context structures, restricting denoising to predetermined token subsets. This creates a mismatch between training and inference, where models must operate over arbitrary configurations, leading to degradation off the training grid. We propose Adaptive Block Diffusion (ABD), which resolves this mismatch by optimizing denoising risk over a distribution of prefix-window configurations. By treating the configuration as a stochastic variable, ABD trains a single model over the full configuration space without architectural changes. We show that generalization across decoding strategies is governed by the support of the training distribution, and that ABD guarantees denoising optimality for any inference policy whose configurations are covered during training. Empirically, ABD exhibits structural invariance across decoding scales, avoiding off-grid collapse and recovering a monotonic relationship between block size and perplexity, while matching or outperforming fixed-block specialists at their target scales.

Gagan Jain• 2026

Related benchmarks

TaskDatasetResultRank
Language modellingLM1B (test)
Perplexity24.76
206
Language ModelingPennTreeBank (PTB)
PPL90.08
157
Language ModelingWikiText
Wikitext PPL29.69
151
Language ModelingLAMBADA
Perplexity (Lambada)51
78
Language ModelingLM1B
Perplexity58.34
65
Language ModelingarXiv
Perplexity37.11
64
Language ModelingAG-News
PPL61.25
45
Language ModelingOpenWebText (OWT) (test)
Perplexity19.57
22
Generative PerplexityOWT L=1024
Generative Perplexity25.22
6
Generative PerplexityOWT L=2048
Generative Perplexity24.46
5
Showing 10 of 10 rows

Other info

Follow for update