FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing
About
Text-to-audio (TTA) generation has made significant strides, yet achieving precise and consistent audio editing remains a major challenge. However, existing methods struggle to balance temporal consistency with background preservation. In this paper, we propose FreeSonic, a training-free framework leveraging the state-of-the-art Rectified Flow-based TangoFlux model. FreeSonic utilizes an optimized inversion-reverse process and joint text-audio attention maps for precise target segment extraction. For content editing, a novel scheduled attention decoupling confines modifications to target regions while preserving original acoustic context. Furthermore, task-oriented noise injection enhances versatility for tasks such as audio removal and non-rigid replacement. Extensive experimental results demonstrate that FreeSonic achieves a superior balance by providing a high-fidelity and efficient solution for precise and consistent audio editing. Project and demos: https://free-sonic.github.io/
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Audio Editing | Audio Editing Remove | CLAP Score42 | 6 | |
| Audio Editing | Audio Editing Replace | CLAP Score0.424 | 6 | |
| Audio Editing | Audio Editing Add | CLAP Score37.4 | 6 | |
| Audio Editing | Audio Editing Subjective Study (Add) | Quality3.84 | 5 | |
| Audio Editing | Audio Editing Subjective Study Remove | Quality Score3.71 | 5 | |
| Audio Editing | Audio Editing Subjective Study Replace | Quality3.67 | 5 |