Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

About

Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiTs) remain insufficiently understood. In this work, we systematically analyze MAs in representative DiTs and find that they are spatially distributed across image tokens while concentrated in a small set of fixed feature dimensions. We further show that these dimensions are closely aligned with AdaLN residual scaling factors and are primarily modulated by the denoising timestep rather than text conditions. This structure leads to two task-dependent effects: for generation, MAs are critical for fine-grained detail synthesis while having limited influence on global semantics; for understanding, their shared high-magnitude directions make raw DiT features overly similar across spatial tokens and weaken dense feature discrimination. Based on these findings, we introduce Eliciting Massive Activation (EMA), a training-free framework that leverages Massive Activations (MAs) as a unified modulation signal to improve both generative and representational capabilities of DiTs. For generation, EMA proposes MA-driven Detail G}uidance (DG), which suppresses MA dimensions to construct a detail-deficient counterfactual prediction and guides sampling toward finer visual details. DG further supports efficient partial-forward inference, integration with classifier-free guidance, and token-level Local DG for refining selected image regions. For understanding, EMA introduces MA-modulated REPresentation extraction (MREP), which uses pretrained AdaLN channel-wise modulation to reduce MA directional dominance and concatenates spatially normalized MA maps to preserve useful spatial structure. Extensive experiments demonstrate that EMA consistently improves both the generation quality and representation capability of DiTs.

Chaofan Gan, Zicheng Zhao, Yuanpeng Tu, Xi Chen, Ziran Qin, Tieyuan Chen, Supavadee Aramvith, Mehrtash Harandi, Weiyao Lin• 2026

Related benchmarks

TaskDatasetResultRank
Class-conditional Image GenerationImageNet 256x256
Inception Score (IS)179.3
1021
Text-to-Image GenerationHPS v2.1
Overall Score31.57
153
Semantic CorrespondencePF-PASCAL
PCK @ alpha=0.196.1
116
Video GenerationVBench (test)
Semantic Score75.06
82
Class-unconditional image generationImageNet 256x256
FID9.68
47
Semantic CorrespondenceSPair-71k
PCK @ 0.018.6
40
Semantic CorrespondenceAP-10K Intra-species (test)
PCK@0.1063
29
Depth EstimationNYU V2
RMSE0.22
22
Semantic CorrespondenceAP-10K cross-family
PCK@0.1049.6
21
Text-to-Image GenerationPick-a-Pic
Clipscore28.12
16
Showing 10 of 14 rows

Other info

Follow for update