Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

About

We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis. Inspired by traditional Foley workflows, our approach enables incremental generation of complementary sounds, allowing users to author multiple sound events induced by a video. To avoid the need for costly multi-reference video-audio datasets, each generation step is formulated as a negatively guided V2A process that discourages duplication of sounds already present in previously generated tracks. The guidance model is trained by finetuning a pre-trained V2A model on audio pairs from non-overlapping segments of the same video, encouraging it to leverage acoustic context while remaining visually grounded, and enabling training with standard single-reference audiovisual datasets. Objective and subjective evaluations demonstrate that our method enhances the separability of generated sounds at each step and improves the overall quality of the final composite audio, outperforming existing baselines. Our project page is available at: https://ahykw.github.io/sbsv2a/.

Akio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji• 2025

Related benchmarks

TaskDatasetResultRank
Composite Audio SynthesisMulti-Caps VGGSound (test)
FD (PANNs)6.47
9
Audio SynthesisMovieGen-Audio-Bench
IS6.63
5
Audio GenerationVGGSound Multi-Caps
Separability3.35
2
Composite Audio GenerationAudioCaps (test)
IS7.08
2
Audio GenerationMovie Gen Audio Bench Individual Audio Tracks
CLAP A-A66.32
2
Individual Audio Track GenerationAudioCaps (test)
CLAP A-A60.43
2
Video-to-Audio GenerationMulti-Caps VGGSound (test)
Audio Quality71.36
1
Showing 7 of 7 rows

Other info

Follow for update