Efficient Diffusion Training via Min-SNR Weighting Strategy

About

Denoising diffusion models have been a mainstream approach for image generation, however, training these models often suffers from slow convergence. In this paper, we discovered that the slow convergence is partly due to conflicting optimization directions between timesteps. To address this issue, we treat the diffusion training as a multi-task learning problem, and introduce a simple yet effective approach referred to as Min-SNR-$\gamma$. This method adapts loss weights of timesteps based on clamped signal-to-noise ratios, which effectively balances the conflicts among timesteps. Our results demonstrate a significant improvement in converging speed, 3.4$\times$ faster than previous weighting strategies. It is also more effective, achieving a new record FID score of 2.06 on the ImageNet $256\times256$ benchmark using smaller architectures than that employed in previous state-of-the-art. The code is available at https://github.com/TiankaiHang/Min-SNR-Diffusion-Training.

Tiankai Hang, Shuyang Gu, Chen Li, Jianmin Bao, Dong Chen, Han Hu, Xin Geng, Baining Guo• 2023

Related benchmarks

Task	Dataset	Result
Class-conditional Image Generation	ImageNet 256x256	--	967
Unconditional Image Generation	CIFAR-10 (test)	FID5.77	223
Class-conditional Image Generation	ImageNet 64x64	FID2.28	170
Text-to-Image Generation	MS-COCO	FID13.92	145
Text-to-Image Generation	PartiPrompts	CLIP Score29.87	26
Unconditional Image Generation	LSUN Church (test)	FID10.82	17
Image Synthesis	ImageNet 256x256	FID1.57	16
Unconditional Image Generation	LSUN Bedroom (test)	FID6.41	14
Text-to-Image Generation	ImageNet	FID27.59	9
Image Generation	CelebA 64x64 50k samples	FID1.6	7

Showing 10 of 10 rows

Other info

Code

Follow for update

@wizwand_team Discord