DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast
About
Text-guided audio editing aims to modify the language-specified acoustic content while preserving edit-irrelevant source components. Existing training-free methods typically rely on inversion-based editing. While inversion-free editing is appealing as it decreases computational overhead and reconstruction errors, it remains largely unexplored for audio editing. The key challenge is to construct a source-to-target editing path through diffusion denoising dynamics. In this paper, we introduce DirectAudioEdit, the first attempt to develop a training-free and inversion-free method for audio editing. Experiments on music and event-level benchmarks across two backbones show that DirectAudioEdit reduces macro-averaged FAD and KL by 15.9% and 15.8% compared with DDPM inversion, while achieving up to 64.5% editing speedup.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Audio Event Removal | Event-level Editing Benchmark | CLAP43.87 | 8 | |
| Audio Event Replacement | Event-level Editing Benchmark | CLAP Score45.82 | 8 | |
| Audio Event Addition | Event-level Editing Benchmark | CLAP Score47.65 | 8 | |
| Music Editing | Music Editing Benchmark | CLAP37.4 | 8 | |
| Audio Editing | Audio Editing Evaluation Set Replacement task | MOS3.43 | 4 |