SyncTweedies: A General Generative Framework Based on Synchronized Diffusions
About
We introduce a general framework for generating diverse visual content, including ambiguous images, panorama images, mesh textures, and Gaussian splat textures, by synchronizing multiple diffusion processes. We present exhaustive investigation into all possible scenarios for synchronizing multiple diffusion processes through a canonical space and analyze their characteristics across applications. In doing so, we reveal a previously unexplored case: averaging the outputs of Tweedie's formula while conducting denoising in multiple instance spaces. This case also provides the best quality with the widest applicability to downstream tasks. We name this case SyncTweedies. In our experiments generating visual content aforementioned, we demonstrate the superior quality of generation by SyncTweedies compared to other synchronization methods, optimization-based and iterative-update-based methods.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| 3D mesh texturing | Objaverse (test) | FID156.8 | 7 | |
| Text-guided 3D mesh texturing | Objaverse 350 (mesh, prompt) pairs | FID163.1 | 6 | |
| Visual Anagram Synthesis | LLM-Augmented dataset | Amin25.38 | 6 | |
| Visual Anagram Synthesis | CIFAR10 2-views | Amin0.2568 | 6 | |
| Wide image generation | 15 text prompts from prior works | Intra-LPIPS0.62 | 5 | |
| 3D mesh texturing | 3D mesh texturing (test) | KID186.6 | 4 | |
| Optical illusion generation | Optical illusion generation | FID255.4 | 4 | |
| Wide image generation | Wide image generation | FID85.95 | 4 | |
| Ambiguous Image Generation | DeepFloyd-IF | KID215.1 | 4 | |
| Mask-based Text-to-Image Generation | Mask-based T2I generation | KID117.4 | 4 |