Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion
About
Forecasting the evolution of dynamic environments is crucial for autonomous agents. While generative world models have recently achieved high photorealism in 2D video synthesis by mixing ego-motion and environmental dynamics within the image plane, they exhibit physical inconsistencies, such as morphing or vanishing objects, especially over long time horizons. In this paper, we propose FR3D, a world model that predicts a persistent 3D latent representation for future dynamic 3D reconstruction. Unlike prior works that treat the world as a sequence of image-based features, FR3D explicitly decouples the 3D evolution of the scene from the agent's trajectory, treating the inferred ego-motion as a latent proxy for action. This disentanglement resolves the ambiguities between self-motion and world-motion, ensuring geometric consistency into the future. Furthermore, we introduce a teacher-student distillation strategy that leverages the spatial "common sense" of off-the-shelf foundation models, leading to robust zero-shot generalization. Extensive experiments demonstrate FR3D's strong performance for future dynamic 3D reconstruction from monocular observations across multiple datasets, even 2 seconds into the future. Project page: https://fr3d-wm.github.io.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Depth Prediction | Dynamic-RE10K | AbsR0.02 | 6 | |
| Depth Estimation | nuScenes (val) | AbsR (T+0.75s)0.182 | 5 | |
| Depth Forecasting | KITTI | AbsR (T+0.6s)0.122 | 5 | |
| Depth Forecasting | nuScenes | Abs Rel Error (T+0.75s)17.6 | 5 | |
| Pose Estimation | KITTI T+1.0s | ATE0.256 | 3 | |
| Pose Estimation | KITTI T+2.0s | ATE0.403 | 3 | |
| Pose Estimation | nuScenes T+1.25s | ATE0.192 | 3 | |
| Pose Estimation | nuScenes T+2.5s | ATE0.437 | 3 | |
| Trajectory Forecasting | RE10K Dynamic | ATE0.01 | 2 |