Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
About
We explore novel-view synthesis for dynamic scenes from monocular videos. Prior approaches rely on costly test-time optimization of 4D representations or do not preserve scene geometry when trained in a feed-forward manner. Our approach is based on three key insights: (1) covisible pixels (that are visible in both the input and target views) can be rendered by first reconstructing the dynamic 3D scene and rendering the reconstruction from the novel-views and (2) hidden pixels in novel views can be "inpainted" with feed-forward 2D video diffusion models. Notably, our video inpainting diffusion model (CogNVS) can be self-supervised from 2D videos, allowing us to train it on a large corpus of in-the-wild videos. This in turn allows for (3) CogNVS to be applied zero-shot to novel test videos via test-time finetuning. We empirically verify that CogNVS outperforms almost all prior art for novel-view synthesis of dynamic scenes from monocular videos.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Dynamic View Synthesis | DyCheck iPhone Masked | mPSNR18.63 | 13 | |
| Single-view Novel View Synthesis | CO3D, Objectron, 360°, OmniObject3D, SynView-X Mean (test) | PSNR8.08 | 9 | |
| Single-view video generation | CO3D, Objectron, 360°, OmniObject3D, SynView-X (test) | Temporal Flickering Score96.1 | 9 | |
| Dynamic View Synthesis | DyCheck iPhone Unseen | uPSNR14.97 | 8 | |
| Dynamic View Synthesis | Dycheck iPhone | PSNR16.94 | 8 | |
| Camera-controlled Video Generation | iPhone dataset 15 (test) | PSNR10.105 | 8 | |
| Camera-controlled Video Generation | DAVIS | Subject Consistency81.1 | 8 | |
| Narrow Dynamic View Synthesis | DyCheck iPhone 1.0 (test) | PSNR16.94 | 7 | |
| Narrow Dynamic View Synthesis | Kubric-4D gradual 1.0 (test) | PSNR22.63 | 7 | |
| Novel View Synthesis | Droid, BridgeData V2, and RoboCoin (test) | PSNR11.88 | 7 |