Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Learning Stable Canonical Worlds for Novel View Synthesis and Beyond

About

Feed-forward Gaussian splatting (FFGS) facilitates real-time novel view synthesis, yet current methods often remain tied to view-dependent predictions. As more input views are added, they may accumulate noisy or redundant evidence instead of converging to a stable scene representation. In this paper, we introduce CanonicalGS, a feed-forward pipeline that maps cluttered multi-view observations into a stable, scene-centric representation. CanonicalGS first extracts view-centric evidence from depth, semantic features, and uncertainty estimates, and then aggregates this evidence in a canonical latent world using uncertainty-aware fusion. By emphasizing reliable observations while suppressing uncertain or redundant ones, CanonicalGS produces representations that scale more effectively for novel view synthesis and transfer to downstream visual perception tasks. Experiments show up to a $2.5$ dB improvement in peak signal-to-noise ratio for synthesizing novel views and an $11\%$ gain in semantic segmentation accuracy.

Xiaoyu Xu, Jian Zou, Sheyang Tang, Zhihua Wang, Jing Liao, Kede Ma• 2026

Related benchmarks

TaskDatasetResultRank
Novel View SynthesisRE10K
SSIM86.1
345
Novel View SynthesisDL3DV (test)
PSNR20.21
120
Novel View SynthesisRE10K bounded-view setting
PSNR27.36
7
Novel View SynthesisACID DepthSplat
PSNR28.47
7
Showing 4 of 4 rows

Other info

Follow for update