Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Discrete Copula Diffusion

About

Discrete diffusion models have recently shown significant progress in modeling complex data, such as natural languages and DNA sequences. However, unlike diffusion models for continuous data, which can generate high-quality samples in just a few denoising steps, modern discrete diffusion models still require hundreds or even thousands of denoising steps to perform well. In this paper, we identify a fundamental limitation that prevents discrete diffusion models from achieving strong performance with fewer steps -- they fail to capture dependencies between output variables at each denoising step. To address this issue, we provide a formal explanation and introduce a general approach to supplement the missing dependency information by incorporating another deep generative model, termed the copula model. Our method does not require fine-tuning either the diffusion model or the copula model, yet it enables high-quality sample generation with significantly fewer denoising steps. When we apply this approach to autoregressive copula models, the combined model outperforms both models individually in unconditional and conditional text generation. Specifically, the hybrid model achieves better (un)conditional text generation using 8 to 32 times fewer denoising steps than the diffusion model alone. In addition to presenting an effective discrete diffusion generation algorithm, this paper emphasizes the importance of modeling inter-variable dependencies in discrete diffusion.

Anji Liu, Oliver Broadrick, Mathias Niepert, Guy Van den Broeck• 2024

Related benchmarks

TaskDatasetResultRank
Code GenerationHumanEval (test)--
701
Code GenerationMBPP (test)--
411
Language ModelingOpenWebText
Perplexity23.33
190
Text GenerationOpenWebText
Perplexity176.7
187
Language ModelingLAMBADA zero-shot (test)--
44
Language GenerationLanguage Modeling Evaluation Set
Generative Perplexity (Llama2)22.2
42
Language ModelingPTB zero-shot
Perplexity95.69
35
Language ModelingLM1B zero-shot
Perplexity66.51
30
Language ModelingPubmed zero-shot
Perplexity47.65
30
Language ModelingAG News zero-shot
Perplexity61.35
22
Showing 10 of 15 rows

Other info

Follow for update