Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model

About

We report Zero123++, an image-conditioned diffusion model for generating 3D-consistent multi-view images from a single input view. To take full advantage of pretrained 2D generative priors, we develop various conditioning and training schemes to minimize the effort of finetuning from off-the-shelf image diffusion models such as Stable Diffusion. Zero123++ excels in producing high-quality, consistent multi-view images from a single image, overcoming common issues like texture degradation and geometric misalignment. Furthermore, we showcase the feasibility of training a ControlNet on Zero123++ for enhanced control over the generation process. The code is available at https://github.com/SUDO-AI-3D/zero123plus.

Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, Hao Su• 2023

Related benchmarks

TaskDatasetResultRank
Novel View SynthesisObjaverse
PSNR14.22
17
3D Scene SynthesisSRN Cars
PSNR17.52
11
Single-image normal estimationSingle-image normal estimation efficiency evaluation (test)
Params (M)52.7
10
Surface Normal EstimationSurface Normal Estimation Benchmark
MAE16.8
10
Multi-view Generation3D-FUTURE
PSNR23.5001
9
Multi-view GenerationGSO
PSNR19.6373
9
Novel View SynthesisSketchFab-Cars
PSNR16.73
9
3D Object GenerationGSO 8 (test)
PSNR15.787
7
Multi-view consistencyDreamFusion 414 text prompts (test)
Avg MRC7
7
Multi-View ReconstructionDreamFusion (test)
Avg MRC0.07
7
Showing 10 of 11 rows

Other info

Follow for update