Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

About

Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models excel at global coherence and visual fidelity but suffer from slow inference, while autoregressive models offer efficient and streaming generation at the cost of long-range consistency and exposure bias. We introduce Flex-Forcing, a unified training and inference framework that enables a video diffusion model to seamlessly operate under both bidirectional and autoregressive generation regimes. The core idea is a flexible chunking mechanism jointly defined over the temporal axis and denoising steps. This design allows the model to (1) perform flexible chunking according to different device budgets, (2) perform bidirectional inference across chunks for global structure planning, while generating frames autoregressively within each chunk for efficient and fine-grained synthesis, and (3) perform any-order, any-timestep autoregressive generation without the strict causal constraint. Extensive experiments on multiple video generation benchmarks demonstrate that Flex-Forcing achieves consistently better video quality, long-video stability than strong baselines with a rigid inference schedule, while offering faster inference.

Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat, Weili Nie, Xinchao Wang• 2026

Related benchmarks

TaskDatasetResultRank
Video GenerationVBench 5s
Quality Score86.33
97
Text-to-Video GenerationVBench 5s videos
Total Score85.13
11
Long Video GenerationVBench-Long 30-second videos
Subject Fidelity Score98
4
Showing 3 of 3 rows

Other info

Follow for update