Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

DisCo: World Models with Discrete Camera Motion Control

About

Controllable video world models target interactive world exploration, where models must faithfully execute explicit action commands while preserving visual quality and temporal coherence. However, most existing approaches rely on continuous camera trajectories as action conditions, which often lead to unreliable action following, especially under complex motion sequences. In this work, we identify action representation entanglement as a key bottleneck in controllable video generation, and show that continuous camera representations lead to high feature similarity across distinct motion patterns, degrading action controllability. Based on this insight, we propose DisCo, a controllable video world model that conditions generation on a compact set of discrete action primitives to improve action separability. We further introduce DisCoBench, a comprehensive benchmark for evaluating the ability of models in short-term, long-horizon, and highly dynamic exploration scenarios. Extensive experiments demonstrate that DisCo achieves significantly more reliable action following while preserving visual quality.

Hongrui Huang, Junke Wang, Quanhao Li, Yu-Gang Jiang, Zuxuan Wu• 2026

Related benchmarks

TaskDatasetResultRank
dynamic exploration video generationDisCo-DynamicBench (test)
FVD220.5
6
Exploration video generationDisCo-ActionBench
Subject Consistency88.69
6
Video GenerationDisCo-LongBench
Subject Consistency0.8138
3
Human Preference EvaluationHuman evaluation
Action Following65.5
2
Video GenerationSekai unseen real-world
FVD333.9
2
Showing 5 of 5 rows

Other info

Follow for update