DiCache: Let Diffusion Model Determine Its Own Cache

About

Recent years have witnessed the rapid development of acceleration techniques for diffusion models, especially caching-based acceleration methods. These studies seek to answer two fundamental questions: "When to cache" and "How to use cache", typically relying on predefined empirical laws or dataset-level priors to determine caching timings and adopting handcrafted rules for multi-step cache utilization. However, given the highly dynamic nature of the diffusion process, they often exhibit limited generalizability and fail to cope with diverse samples. In this paper, a strong sample-specific correlation is revealed between the variation patterns of the shallow-layer feature differences in the diffusion model and those of deep-layer features. Moreover, we have observed that the features from different model layers form similar trajectories. Based on these observations, we present DiCache, a novel training-free adaptive caching strategy for accelerating diffusion models at runtime, answering both when and how to cache within a unified framework. Specifically, DiCache is composed of two principal components: (1) Online Probe Profiling Scheme leverages a shallow-layer online probe to obtain an on-the-fly indicator for the caching error in real time, enabling the model to dynamically customize the caching schedule for each sample. (2) Dynamic Cache Trajectory Alignment adaptively approximates the deep-layer feature output from multi-step historical caches based on the shallow-layer feature trajectory, facilitating higher visual quality. Extensive experiments validate DiCache's capability in achieving higher efficiency and improved fidelity over state-of-the-art approaches on various leading diffusion models including WAN 2.1, HunyuanVideo and Flux.

Jiazi Bu, Pengyang Ling, Yujie Zhou, Yibin Wang, Yuhang Zang, Dahua Lin, Jiaqi Wang• 2025

Related benchmarks

Task	Dataset	Result
Text-to-Image Generation	MJHQ-30K	Overall FID20.7	239
Text-to-Image Generation	MS-COCO (val)	--	202
Text-to-Image Generation	PartiPrompts	ImageReward0.97	92
Text-to-Image Generation	MS-COCO (30K)	FID (30K)26.18	72
Text-to-Image Generation	MS COCO 2017	FID28.19	41
Text-to-Image Generation	Image Reward (Calibration)	Image Reward0.61	32
Text-to-Video Generation	HunyuanVideo	LPIPS0.382	30
Image2World generation	PAI-Bench	Domain Score Average0.855	17
Text2World (T2W) Generation	PAI-Bench	Latency (s)40.82	14
Text-to-Image Generation	FLUX.1 1.0 (dev)	FLOPs (T)958.1	12

Showing 10 of 13 rows

Other info

Follow for update

@wizwand_team Discord