Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

About

Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow, namely RGB-DF, provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to the low-level end-effector actions demanded by robotic systems, thereby narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces temporally and spatially coherent 4D predictions, and that RynnWorld-4D-Policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.

Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li• 2026

Related benchmarks

TaskDatasetResultRank
Stack bowlsReal-world Robotic Tasks
Success Rate65.71
12
Visual Synthesis4D World Modeling
IQ0.635
12
Depth Estimation4D World Modeling
AbsRel0.31
9
Bimanual LiftingRobotic Manipulation Tasks
Success Rate97.14
8
Hand OverRobotic Manipulation Tasks
Success Rate28.57
8
Lid PlacementRobotic Manipulation Tasks
Success Rate65.71
8
Block PushingRobotic Manipulation Tasks
Success Rate97.14
8
Dual PickingRobotic Manipulation Tasks
Success Rate94.29
8
Optical Flow Estimation4D World Modeling
AEPE0.17
6
Showing 9 of 9 rows

Other info

GitHub

Follow for update