Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling

About

Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and content creation. Recent methods achieve outstanding performance by leveraging camera-controlled video diffusion models, but rely on iterative diffusion sampling, which greatly limits their practical deployment. We observe that geometric forward warping alone can cover the majority of a target view directly from the input image, with only a compact residual left for the encoder to correct. Motivated by this observation, we propose PRISM, a feed-forward framework that decomposes multi-view latent prediction into a parameter-free geometric prior and a learned residual correction, with no diffusion sampling required at inference. To enable generalization from purely synthetic training data, we devise a two-stage training strategy combining latents supervised distillation for geometric generalization and perceptual fine-tuning for appearance quality optimization. Extensive experiments on three benchmarks demonstrate that PRISM achieves competitive reconstruction quality compared with diffusion-based methods, while reducing inference time dramatically to only 36 seconds per scene.

Zhijie Zheng, Xinhao Xiang, Jiawei Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Novel View SynthesisRealEstate10K
PSNR20.43
212
Novel View SynthesisDL3DV
PSNR19.46
92
Showing 2 of 2 rows

Other info

Follow for update