Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

About

Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches typically treat diffusion-based controllable generation and dense prediction as separate tasks, overlooking the potential benefits of jointly modeling the heterogeneous distributions. In this work, we introduce UniGP, a framework built upon MMDiT, which unifies controllable generation and dense prediction through simple joint training, without the need for complex task-specific designs or losses, while preserving the backbone's versatile priors. By learning controllable generation and prediction under different conditions, our model effectively captures the joint distribution of image-geometry pairs. UniGP is capable of versatile controllable generation, dense prediction, and joint generation. Specifically, the proposed UniGP consists of DUGP and a unified dataset training strategy. The former, following the principle of Occam's razor, uses only a copied image branch of MMDiT to model dense distributions beyond RGB, while the latter integrates heterogeneous datasets into a unified training framework to jointly model generation and perception tasks. Extensive experiments demonstrate that our unified model surpasses prior unified approaches and performs on par with specialized methods. Furthermore, we demonstrate that multi-task joint training provides complementary benefits: generative priors enrich perceptual details, while perceptual learning improves structural alignment in generation.

Qin Guo, Hao Luo, Dongxu Yue, Weixuan Jin, Xiao Fu, Fan Wang, Dan Xu• 2026

Related benchmarks

TaskDatasetResultRank
Surface Normal EstimationNYU V2
Mean Angular Error16.4
96
Surface Normal EstimationiBIMS-1
MAE17.3
67
Surface Normal EstimationScanNet Indoor
Mean Error14.9
18
Surface Normal EstimationSintel Outdoor
Accuracy (11.25° Threshold)20.1
14
Zero-shot affine-invariant depth estimationNYU Indoor v2
AbsRel5.2
12
Zero-shot affine-invariant depth estimationScanNet Indoor
AbsRel5.5
12
Zero-shot affine-invariant depth estimationETH3D Various
AbsRel6
9
Zero-shot affine-invariant depth estimationKITTI Outdoor
AbsRel8.3
9
Surface Normal EstimationAverage across benchmarks
Average Rank1.58
8
Depth-conditioned Image GenerationCOCO 5k
FID17.95
7
Showing 10 of 15 rows

Other info

Follow for update