Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards

About

Diffusion-based large language models offer a non-autoregressive alternative for text generation, but enabling them to perform complex reasoning remains challenging. Reinforcement learning has recently emerged as an effective post-training strategy for improving their performance; however, existing methods rely primarily on outcome-based rewards, which provide no direct supervision over the denoising process and often result in poorly structured reasoning that is difficult to interpret and inconsistently supports the final prediction. To address this limitation, we introduce \emph{denoising process reward}, a process-level reinforcement signal defined over the denoising trajectory of diffusion language models. This reward is obtained by estimating the contribution of intermediate denoising intervals to the final task outcome, encouraging the model to favor reasoning trajectories that consistently guide generation toward correct predictions. We further propose an efficient stochastic estimator that reuses standard training rollouts, enabling practical process-level supervision at scale. Experiments on challenging reasoning benchmarks demonstrate that our approach yields consistent improvements in reasoning stability, interpretability, and overall task performance.

Shaoan Xie, Lingjing Kong, Xiangchen Song, Xinshuai Dong, Guangyi Chen, Eric P.Xing, Kun Zhang• 2025

Related benchmarks

Task	Dataset	Result
Mathematical Reasoning	GSM8K	Accuracy (Acc)82.1	337
Mathematical Reasoning	Countdown	Accuracy56.3	252
Logical reasoning	Sudoku	Accuracy22.4	142
Planning	Sudoku	Accuracy20.3	129
Planning	Countdown	Accuracy56.3	89
Commonsense Reasoning	ARC	Accuracy93	61
Advanced Mathematical Reasoning	Math500 256 tokens	Pass@1 Accuracy38.6	15
Grade School Math Word Problems	GSM8k 256 tokens	Pass@180.6	15
Grade School Math Word Problems	GSM8k 512 tokens	Pass@182.1	15
Arithmetic Reasoning	Countdown 512 tokens	Pass@156.3	15

Showing 10 of 14 rows

Other info

Follow for update

@wizwand_team Discord