Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models

About

Flow-matching transformers achieve strong audio separation, yet their attention dynamics are opaque. We adapt established causal-intervention principles into a deterministic, inference-time probing protocol for SAM Audio. Orthogonal probing uncovers a dual-pathway text-conditioning mechanism: additive injections control semantic identity, while cross-attention refines acoustic structure. We observe an asynchronous layerwise convergence: stable layers build temporal scaffolds early, whereas fast layers continue resolving artifacts during sampling. The model also attenuates temporal segmentation cues to maintain continuous-flow stability. Using these insights, we propose Layer-Selective Attention Caching (LSAC), a training-free acceleration method that caches attention in stable layers. Across acoustic complexities, LSAC cuts self-attention computation by about ~25% with negligible quality loss and yields up to 6.7x higher quality retention than naive step reduction.

Yuxuan Chen, Haoyuan Yu, Peize He• 2026

Related benchmarks

TaskDatasetResultRank
Audio SeparationNoisy Tier
|ΔSI-SNR| (dB)0.1
9
Audio SeparationClean Tier
|ΔSI-SNR| (dB)2.5
9
Audio SeparationEnvironmental Tier (Env)
Delta SI-SNR (dB)0.3
9
Showing 3 of 3 rows

Other info

Follow for update