CacheQuant: Comprehensively Accelerated Diffusion Models

About

Diffusion models have gradually gained prominence in the field of image synthesis, showcasing remarkable generative capabilities. Nevertheless, the slow inference and complex networks, resulting from redundancy at both temporal and structural levels, hinder their low-latency applications in real-world scenarios. Current acceleration methods for diffusion models focus separately on temporal and structural levels. However, independent optimization at each level to further push the acceleration limits results in significant performance degradation. On the other hand, integrating optimizations at both levels can compound the acceleration effects. Unfortunately, we find that the optimizations at these two levels are not entirely orthogonal. Performing separate optimizations and then simply integrating them results in unsatisfactory performance. To tackle this issue, we propose CacheQuant, a novel training-free paradigm that comprehensively accelerates diffusion models by jointly optimizing model caching and quantization techniques. Specifically, we employ a dynamic programming approach to determine the optimal cache schedule, in which the properties of caching and quantization are carefully considered to minimize errors. Additionally, we propose decoupled error correction to further mitigate the coupled and accumulated errors step by step. Experimental results show that CacheQuant achieves a 5.18 speedup and 4 compression for Stable Diffusion on MS-COCO, with only a 0.02 loss in CLIP score. Our code are open-sourced: https://github.com/BienLuky/CacheQuant .

Xuewen Liu, Zhikai Li, Qingyi Gu• 2025

Related benchmarks

Task	Dataset	Result
Class-conditional Image Generation	ImageNet 256x256 (val)	Inception Score (IS)213.1	535
Class-conditional Image Generation	ImageNet 256x256 (test)	FID4.03	223
Text-to-Image Generation	MS-COCO	FID23.23	193
Unconditional Image Generation	CIFAR-10 32x32 (test)	FID4.61	137
Image Generation	LSUN Bedroom 256x256 (test)	FID8.85	81
Unconditional Image Generation	LSUN Churches 256 x 256 (test)	FID3.52	18
Text-conditional generation	PartiPrompts	Generation Speed (x)5.2	9

Showing 7 of 7 rows

Other info

Follow for update

@wizwand_team Discord