Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling

About

We introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps). JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense modality branch and is densely connected with the RGB branch. The RGB branch is locked during network fine-tuning, which enables efficient learning of the new modality distribution while maintaining the strong generalization ability of the large-scale pre-trained diffusion model. We demonstrate the effectiveness of JointNet by using RGBD diffusion as an example and through extensive experiments, showcasing its applicability in a variety of applications, including joint RGBD generation, dense depth prediction, depth-conditioned image generation, and coherent tile-based 3D panorama generation.

Jingyang Zhang, Shiwei Li, Yuanxun Lu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Yao Yao• 2023

Related benchmarks

TaskDatasetResultRank
Monocular Depth EstimationNYU V2
Delta 1 Acc86.7
192
Monocular Depth EstimationETH3D
AbsRel12.63
173
Monocular Depth EstimationDIODE
AbsRel20.02
161
Monocular Depth EstimationScanNet
AbsRel14.81
111
Monocular Depth EstimationKITTI
AbsRel13.74
92
Zero-shot affine-invariant depth estimationScanNet Indoor
AbsRel11.9
12
Zero-shot affine-invariant depth estimationNYU Indoor v2
AbsRel13.6
12
Zero-shot affine-invariant depth estimationKITTI Outdoor
AbsRel29.9
9
Zero-shot affine-invariant depth estimationETH3D Various
AbsRel19.2
9
Depth-conditioned Image GenerationCOCO 5k
FID25.66
7
Showing 10 of 11 rows

Other info

Follow for update