Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation

About

Recent unified audio generation models can support diverse tasks across speech, sound effects, and music, but most of them still focus on isolated task-level synthesis. However, real video production often requires multiple components of a complete audio track to be generated jointly and consistently for the same video. We present Foley-Omni, a unified multimodal audio generation model that extends isolated task-level synthesis to complete video soundtrack generation by jointly modeling speech, sound effects, and music within a shared latent generation process. To support training and reproducible evaluation, we develop an audiovisual data curation pipeline and introduce V2ST-Bench, a benchmark for holistic video soundtrack generation evaluation. Experiments show that Foley-Omni achieves competitive performance with expert systems on individual synthesis tasks, while improving speech intelligibility, audiovisual consistency and perceptual quality for mixed soundtrack generation.

Ye Tao, Lupeng Liu, Xuenan Xu, Jiasun Feng, Jiarui Wang, Ying Qin, Shuiyang Mao, Wei Liu, Shuai Wang• 2026

Related benchmarks

TaskDatasetResultRank
Video-to-Audio GenerationVGGSound
FD_VGG1.57
32
Text-to-MusicDownstream Audio Generation (TTM)
CLAP Score0.374
12
Visual Text-to-SpeechLRS2 zero-shot
WER13
5
Visual Text-to-SpeechGRID (seen-speaker)
WER15.3
5
Text-to-Audio GenerationTTA Text-to-Audio
CLAP Score46
5
Complete video soundtrack generationV2ST-Bench
CLAP Score0.27
4
Text-to-speech generationTTS
WER2.31
4
Showing 7 of 7 rows

Other info

Follow for update